China Mobile's JT-4.1 Flash Challenges DeepSeek With a Free 236B Model

China Mobile's JT-4.1 Flash 236B A21B scores 39 on the AA Intelligence Index, closing the gap to China's AI frontier with a massive leap from its predecessor.

·
·
China Mobile's JT-4.1 Flash Challenges DeepSeek With a Free 236B Model
  • New model: China Mobile released JT-4.1 Flash 236B A21B on July 9, scoring 39 on the Artificial Analysis Intelligence Index.
  • Big jump: Scores 13 points higher than predecessor JT-35B-Flash, with HLE reasoning score tripling from 6% to 16.1%.
  • Free to use: Priced at $0.00 per 1M input and output tokens via China Mobile's first-party API.
  • Banking standout: Scores 28% on τ³-Banking, beating DeepSeek V4 Pro (26%) and GLM-5.2 (27%) on this agentic finance benchmark.
  • Hallucination tradeoff: 43% hallucination rate but abstains more (59% attempt rate vs 74%), cutting hallucinations by 20 points vs predecessor.
  • Telecom AI race: Part of China Mobile's broader push to become an AI platform, with 300+ models on its MoMA platform and $6.83B in Q1 AI revenue.

China Mobile just dropped JT-4.1 Flash 236B A21B, the latest entry in its growing JT model series. The release marks a meaningful step forward for a telecom giant that is betting its next decade on becoming a full-stack AI infrastructure company -- not just a carrier selling bandwidth.

What just shipped

Released on July 9, 2026, JT-4.1 Flash 236B A21B was created by China Mobile. The model is among the leading models in intelligence for its price tier, supports text input and output, and has a 256k token context window. It is not a reasoning model -- it provides direct responses without extended chain-of-thought reasoning.

JT-4.1 Flash 236B A21B scores 39 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models, which average a score of 9. Pricing is $0.00 per 1M input tokens and $0.00 per 1M output tokens. That makes it free to use through China Mobile's first-party API -- a positioning move that matters more than it might first appear.

The numbers that matter

The Artificial Analysis Intelligence Index v4.1 is a composite benchmark that rolls up nine evaluations spanning agentic tasks, coding, scientific reasoning, and knowledge retrieval. Here is how JT-4.1 Flash performed on the key sub-evals:

  • Humanity's Last Exam (HLE): 16.1% -- a significant jump from JT-35B-Flash's 6%. HLE is a benchmark of 2,500 expert-level questions across math, science, and the humanities, designed to measure an LLM's reasoning capabilities, not just its ability to pattern match, evaluating how well a model handles expert-level problems across many academic domains. Even at 16.1%, there is a large gap to GLM-5.2's 40% on the same eval.
  • τ³-Banking: 28% -- an agentic tool-use benchmark focused on banking workflows. This puts JT-4.1 Flash ahead of DeepSeek V4 Pro (26%) and GLM-5.2 (27%), and just 5 points behind the leading model.
  • AA-Omniscience Index: -10.1 -- a knowledge reliability score that rewards correct answers and penalizes hallucinations. A negative score means the model produces more wrong answers than right ones, but the +13.1 point improvement over JT-35B-Flash is notable.

The hallucination story deserves its own paragraph. JT-4.1 Flash has a 43% hallucination rate and a 23% accuracy rate on AA-Omniscience. That sounds alarming, but the model made a deliberate tradeoff: it only attempts to answer 59% of questions (down from 74% for its predecessor), abstaining more often rather than guessing. The result is a 20-point reduction in hallucinations. It is a conservative strategy -- the model is choosing silence over confident wrongness.

On token efficiency, JT-4.1 Flash uses roughly 14k output tokens per Intelligence Index task. When evaluated on the full Intelligence Index, it generated 74M output tokens, which is at the higher end compared to other non-reasoning models in a similar price tier, where the median is 8.6M. That verbosity is unusual for a non-reasoning model and suggests the model is doing more internal work than a typical fast-response system.

Where it sits in the Chinese AI landscape

The Chinese AI frontier is intensely competitive right now. On the Artificial Analysis Intelligence Index v4.1, GLM-5.2 scores 51, Qwen3.7 Max scores 46, and Kimi K2.6 scores 43. JT-4.1 Flash at 39 trails all three, but the gap has closed substantially from where JT-35B-Flash was sitting just two months ago.

The comparison that matters most for China Mobile is not against GLM or Qwen -- it is against where the JT series was. The previous model, JT-35B-Flash, scored around 26 on the same index. A 13-point jump in roughly two months, combined with the move to a much larger 236B parameter architecture, signals that China Mobile's AI team is iterating fast.

Why a telecom company is building frontier models

This release only makes sense in the context of what China Mobile is trying to become. China Mobile has announced plans to accelerate its transformation from a traditional telecommunications operator into a world-class sci-tech service enterprise, with Chairman Chen Zhongyue positioning communications services, computing services, and AI services as the company's three principal business pillars.

China Mobile is pursuing a mass-market approach focused on large-scale token distribution, and officially launched its Token Operations Ecosystem in May 2026, with regional branches rapidly introducing commercial offerings. The platform integrates around 300 AI models, and China Mobile said centralized operations could help reduce per-token costs by around 30%.

The strategic logic is straightforward: the moves mark a shift in how China's telecom sector hopes to profit from generative AI, as operators attempt to transform computing power and AI model access into a utility-like service resembling traditional mobile data packages. Having a proprietary model in that stack -- one that is free to use -- is a way to anchor users to China Mobile's own infrastructure rather than routing them to DeepSeek or Qwen.

According to data from China's National Data Administration, average daily token usage in China has jumped from 100 billion at the start of 2024 to 140 trillion by March 2026 -- an increase of more than 1,000 times in just two years. China Mobile wants to own a meaningful slice of that demand.

The bigger infrastructure play

Beijing announced plans to invest 2 trillion yuan ($295 billion) over five years in AI datacenter infrastructure, with China Mobile and China Telecom designated as the primary operators of a national AI compute network, and Huawei supplying the majority of AI chips. China Mobile says its AI-related revenue was 46.6 billion yuan ($6.83 billion) in Q1 2026, up 12.7% year-on-year.

China Mobile is building a closed-loop AI ecosystem, aggregating AI models on one end while linking them to devices and applications on the other. Its Mobile Model Management Platform (MoMA) connects not only China Mobile's in-house large language model, but also more than 300 mainstream AI models, including DeepSeek, Qwen, Doubao, and GLM. JT-4.1 Flash is the in-house anchor of that platform.

Who wins and who loses

The clearest winner is China Mobile's enterprise and government customer base. A free, capable, domestically-developed model that runs on China Mobile's own infrastructure is a compelling offer for organizations with data sovereignty requirements. The model's strong τ³-Banking score suggests it was specifically tuned for financial services workflows -- a high-value vertical.

The competitive pressure lands on mid-tier Chinese model providers. JT-4.1 Flash at 39 on the Intelligence Index, offered for free, undercuts the value proposition of smaller labs that charge for models in the same capability range. For the true frontier -- GLM-5.2, Qwen3.7 Max, DeepSeek V4 Pro -- the gap is still large enough that JT-4.1 Flash is not a direct threat yet.

The model's limitations are real. JT-4.1 Flash 236B A21B is not currently available through any API providers benchmarked by Artificial Analysis, and no third-party API providers are currently available for the model. That means access is gated entirely through China Mobile's own platform -- which is by design, but limits the model's reach outside China Mobile's existing customer base.

What to watch next

The JT series is clearly on a rapid release cadence. JT-35B-Flash launched in May, JT-4.1 Flash arrived in July. If that pace continues, a reasoning variant or multimodal upgrade could arrive before the end of the year. The model's name -- 236B A21B -- follows the Mixture-of-Experts naming convention (236 billion total parameters, 21 billion active per token), suggesting China Mobile is now operating at a scale comparable to DeepSeek V3's architecture class, even if the benchmark scores have not caught up yet.

For teams building on Chinese AI infrastructure or evaluating models for finance and banking use cases, JT-4.1 Flash is worth a benchmark run. It is free, it has a 256k context window, and its τ³-Banking performance is genuinely competitive with the Chinese frontier. Just go in with clear eyes about the hallucination rate -- and test your specific use case before committing.

Comments

avatar