China Mobile's JT-4.1 Flash Challenges DeepSeek With a Free 236B Model
China Mobile's JT-4.1 Flash 236B A21B scores 39 on the AA Intelligence Index, closing the gap to China's AI frontier with a massive leap from its predecessor.

- New model: China Mobile released JT-4.1 Flash 236B A21B on July 9, scoring 39 on the Artificial Analysis Intelligence Index.
- Big jump: Scores 13 points higher than predecessor JT-35B-Flash, with HLE reasoning score tripling from 6% to 16.1%.
- Free to use: Priced at $0.00 per 1M input and output tokens via China Mobile's first-party API.
- Banking standout: Scores 28% on τ³-Banking, beating DeepSeek V4 Pro (26%) and GLM-5.2 (27%) on this agentic finance benchmark.
- Hallucination tradeoff: 43% hallucination rate but abstains more (59% attempt rate vs 74%), cutting hallucinations by 20 points vs predecessor.
- Telecom AI race: Part of China Mobile's broader push to become an AI platform, with 300+ models on its MoMA platform and $6.83B in Q1 AI revenue.
China Mobile just dropped JT-4.1 Flash 236B A21B, the latest entry in its growing JT model series. The release marks a meaningful step forward for a telecom giant that is betting its next decade on becoming a full-stack AI infrastructure company -- not just a carrier selling bandwidth.
What just shipped
Released on July 9, 2026, JT-4.1 Flash 236B A21B was created by China Mobile. The model is among the leading models in intelligence for its price tier, supports text input and output, and has a 256k token context window. It is not a reasoning model -- it provides direct responses without extended chain-of-thought reasoning.
JT-4.1 Flash 236B A21B scores 39 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models, which average a score of 9. Pricing is $0.00 per 1M input tokens and $0.00 per 1M output tokens. That makes it free to use through China Mobile's first-party API -- a positioning move that matters more than it might first appear.
The numbers that matter
The Artificial Analysis Intelligence Index v4.1 is a composite benchmark that rolls up nine evaluations spanning agentic tasks, coding, scientific reasoning, and knowledge retrieval. Here is how JT-4.1 Flash performed on the key sub-evals:
- Humanity's Last Exam (HLE): 16.1% -- a significant jump from JT-35B-Flash's 6%. HLE is a benchmark of 2,500 expert-level questions across math, science, and the humanities, designed to measure an LLM's reasoning capabilities, not just its ability to pattern match, evaluating how well a model handles expert-level problems across many academic domains. Even at 16.1%, there is a large gap to GLM-5.2's 40% on the same eval.
- τ³-Banking: 28% -- an agentic tool-use benchmark focused on banking workflows. This puts JT-4.1 Flash ahead of DeepSeek V4 Pro (26%) and GLM-5.2 (27%), and just 5 points behind the leading model.
- AA-Omniscience Index: -10.1 -- a knowledge reliability score that rewards correct answers and penalizes hallucinations. A negative score means the model produces more wrong answers than right ones, but the +13.1 point improvement over JT-35B-Flash is notable.