Z.ai's GLM-5.2 Beats GPT-5.5 on Coding With a 1M Context Window
Z.ai's GLM-5.2 is a 753B open-weight MoE model with a true 1M-token context window, beating GPT-5.5 on coding benchmarks under an MIT license
PRO- GLM-5.2 is a 753B open-weight MoE model from Z.ai, available now on Hugging Face under an MIT license with no regional restrictions.
- Context window jumps 5× from 200K to 1 million tokens, with up to 131K output tokens per response, trained specifically on long coding-agent trajectories.
- New IndexShare architecture reduces per-token FLOPs by 2.9× at 1M context by sharing sparse attention indexers across every four transformer layers.
- Beats GPT-5.5 on SWE-bench Pro (62.1 vs 58.6), Terminal-Bench 2.1 (82.7 vs 83.4), and FrontierSWE (74.4 vs 72.6); trails Claude Opus 4.8 on most benchmarks.
- Two thinking modes (High and Max) let you trade token usage for performance; Max uses ~85K output tokens per task, High roughly halves that.
- Trained entirely on Huawei Ascend 910B chips; API pricing starts at $12.60/month, roughly 1/6th the cost of comparable closed models.
GLM-5.2 is the latest flagship from Z.ai (the international brand of Zhipu AI, spun out of Tsinghua University), and it arrives at a charged moment for the open-source AI world. Announced just days after the US government ordered Anthropic to pull Claude Fable 5 and Claude Mythos 5 from international users, Z.ai founder Jie Tang made the subtext explicit: "science should be global." The timing was deliberate, and the model backs up the statement.
GLM-5.2 is a 753-billion-parameter Mixture-of-Experts (MoE) model , meaning it has 753B total weights but only activates roughly 40B parameters per token, keeping inference costs manageable. It is the third major release in Z.ai's GLM-5 series, trained entirely on Huawei Ascend 910B chips using the MindSpore framework , no NVIDIA hardware. The weights are available now on Hugging Face under an MIT license, with no regional restrictions.
A 1M context that actually works
The headline feature is a genuine 1-million-token context window. The window jumps from 200K (GLM-5.1) to 1 million tokens, with outputs capped at 131,072 tokens per response. That five-fold increase is not just a number on a spec sheet. Z.ai explicitly trained the model on long, messy coding-agent trajectories at 1M scale, covering large-scale implementation, automated research, performance optimization, and complex debugging. The goal was reliability under real engineering pressure, not just token acceptance.
The practical consequence: for agentic coding work , long repo ingests, multi-file refactors, marathon sessions , this matters. The previous 200K limit meant agents had to compact frequently. At 1M tokens, a GLM-5.2-backed coding session can hold substantially more context before hitting the wall.
The architecture trick that makes 1M affordable
Running attention over 1 million tokens is brutally expensive. Standard transformers recompute attention indices at every layer, and at extreme context lengths that cost dominates everything else. GLM-5.2 introduces a technique called IndexShare to solve this.
IndexShare reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. To understand why this works: GLM-5 uses a sparse attention mechanism called DSA (Dynamic Sparse Attention), where a lightweight "indexer" network first decides which tokens are relevant before computing full attention. IndexShare's insight is that adjacent transformer layers tend to need the same relevant tokens , so you only run the indexer once every four layers and share the result. Experimental results show that IndexCache can remove 75% of indexer computations with negligible quality degradation, achieving up to 1.82× prefill speedup and 1.48× decode speedup.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.