Alibaba's Qwen3.8-Max Sneaks Onto Leaderboards Before Its 2.4T Official Launch

Alibaba's 2.4T-parameter Qwen3.8-Max arrives with open weights promised, frontier coding claims, and a stealth arena debut that the community unmasked

·
·
AuthorQwen
Read2 min
  • Qwen3.8-Max announced: Alibaba's new flagship at 2.4T parameters (sparse MoE), claiming second place behind Claude Fable 5 globally.
  • Open weights promised next week: Both Qwen3.8-Max and a smaller Qwen3.8-27B will go open-weight, breaking the recent Max-tier closed pattern.
  • Stealth arena debut: The model was caught on Code Arena as anonymous model "kaleb" — identified by its Qwen tokenizer signature — before the official announcement.
  • Autonomous coding demo: The model built the oh-my-cli repo over 16 days with 422 commits, zero tool-call failures in independent testing.
  • Pricing: $2.0/M input, $6.0/M output — try it at Qwen Studio or via the API.
  • No public benchmarks yet: All performance claims are Alibaba's own internal evals; no independent scores from Artificial Analysis, Arena.AI, or Hugging Face exist as of publication.

Alibaba has officially announced Qwen3.8-Max, its largest model to date at a claimed 2.4 trillion parameters, paired with a promise that open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B will drop next week. The model is live today as a preview through Qwen Studio and the Qwen API, with Alibaba positioning it as second only to Claude Fable 5 among all frontier models.

The stealth arena debut

Before the official announcement, something unusual happened on the Code Arena leaderboard. On July 18, an anonymous model called "kaleb" appeared, introducing itself as "Claude," a training artifact left over from Anthropic distillation. Within 24 hours, the community cracked the identity. The tell was a quirk in token generation: the model produced tokens tagged as PostalCodesNL, a pattern unique to Alibaba's Qwen tokenizer. Alibaba confirmed it the next day: "kaleb" was Qwen3.8-Max.

On the Code Arena coding leaderboard, Qwen3.8-Max leads Kimi K3 by approximately 6 Elo points, a meaningful but not dominant margin, despite Kimi K3 having 2.8T parameters versus Qwen3.8-Max's 2.4T.

What Alibaba is actually shipping

Alibaba describes Qwen3.8-Max as a 2.4-trillion-parameter sparse mixture-of-experts system that handles text, images, video, and documents. In a sparse MoE architecture, only a fraction of the network activates per token, which keeps inference costs manageable. Alibaba hasn't disclosed the active-parameter count or the MoE configuration, so 2.4T is a headline number, not a compute figure. For scale, DeepSeek V4 Pro's 1.6T total only activates about 49B parameters per token, roughly 3% of the network.

Specs pulled from Qwen Cloud integration metadata:

  • 983,616-token context window and 131,072-token maximum output

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves