China Mobile's JT-4.1 Flash Challenges DeepSeek With a Free 236B Model

China Mobile's JT-4.1 Flash 236B A21B scores 39 on the AA Intelligence Index, closing the gap to China's AI frontier with a massive leap from its predecessor.

·
·
China Mobile's JT-4.1 Flash Challenges DeepSeek With a Free 236B Model
  • New model: China Mobile released JT-4.1 Flash 236B A21B on July 9, scoring 39 on the Artificial Analysis Intelligence Index.
  • Big jump: Scores 13 points higher than predecessor JT-35B-Flash, with HLE reasoning score tripling from 6% to 16.1%.
  • Free to use: Priced at $0.00 per 1M input and output tokens via China Mobile's first-party API.
  • Banking standout: Scores 28% on τ³-Banking, beating DeepSeek V4 Pro (26%) and GLM-5.2 (27%) on this agentic finance benchmark.
  • Hallucination tradeoff: 43% hallucination rate but abstains more (59% attempt rate vs 74%), cutting hallucinations by 20 points vs predecessor.
  • Telecom AI race: Part of China Mobile's broader push to become an AI platform, with 300+ models on its MoMA platform and $6.83B in Q1 AI revenue.

China Mobile just dropped JT-4.1 Flash 236B A21B, the latest entry in its growing JT model series. The release marks a meaningful step forward for a telecom giant that is betting its next decade on becoming a full-stack AI infrastructure company -- not just a carrier selling bandwidth.

What just shipped

Released on July 9, 2026, JT-4.1 Flash 236B A21B was created by China Mobile. The model is among the leading models in intelligence for its price tier, supports text input and output, and has a 256k token context window. It is not a reasoning model -- it provides direct responses without extended chain-of-thought reasoning.

JT-4.1 Flash 236B A21B scores 39 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models, which average a score of 9. Pricing is $0.00 per 1M input tokens and $0.00 per 1M output tokens. That makes it free to use through China Mobile's first-party API -- a positioning move that matters more than it might first appear.

The numbers that matter

The Artificial Analysis Intelligence Index v4.1 is a composite benchmark that rolls up nine evaluations spanning agentic tasks, coding, scientific reasoning, and knowledge retrieval. Here is how JT-4.1 Flash performed on the key sub-evals:

  • Humanity's Last Exam (HLE): 16.1% -- a significant jump from JT-35B-Flash's 6%. HLE is a benchmark of 2,500 expert-level questions across math, science, and the humanities, designed to measure an LLM's reasoning capabilities, not just its ability to pattern match, evaluating how well a model handles expert-level problems across many academic domains. Even at 16.1%, there is a large gap to GLM-5.2's 40% on the same eval.
  • τ³-Banking: 28% -- an agentic tool-use benchmark focused on banking workflows. This puts JT-4.1 Flash ahead of DeepSeek V4 Pro (26%) and GLM-5.2 (27%), and just 5 points behind the leading model.
  • AA-Omniscience Index: -10.1 -- a knowledge reliability score that rewards correct answers and penalizes hallucinations. A negative score means the model produces more wrong answers than right ones, but the +13.1 point improvement over JT-35B-Flash is notable.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves