Baidu's ERNIE-4.5-21B-A3B-Thinking Beats Dense 70B Models on a Single GPU
Baidu's compact 21B Mixture-of-Experts model activates only 3B parameters per token, targeting deep reasoning, tool use, and 128K context on a single GPU.
PRO- Baidu released ERNIE-4.5-21B-A3B-Thinking, a 21B MoE reasoning model activating only 3B parameters per token.
- Apache 2.0 license, native 128K context, and OpenAI-compatible tool calling out of the box.
- Scores 89.8 on ZebraLogic, 87.77 on BBH, over 90 on HumanEval+, competitive with Gemini 2.5 Pro on many tasks.
- Runs on a single 80GB GPU via FastDeploy, vLLM, or SGLang.
- Long context built by progressive RoPE scaling from 10K to 500K base plus FlashMask attention.
- Trails frontier models on hardest math and scientific QA, and thinking traces are notably longer.
Baidu has quietly released a reasoning-focused open model that punches well above its parameter count. ERNIE-4.5-21B-A3B-Thinking is a Mixture-of-Experts model with 21B total parameters that activates only about 3B per token, ships under Apache 2.0, and carries a 128K context window. The goal is chain-of-thought reasoning depth without the compute cost of a dense frontier model.
What's inside
The architecture is a sparse MoE stack with an aggressive expert count for its size. Per the model card:
- 21B total parameters, 3B activated per token
- 28 layers, 20 query heads and 4 key/value heads (grouped-query attention)
- 64 text experts with 6 activated per token, plus 2 shared experts always on
- 131,072 token context length
The routing logic is where most of the engineering effort shows. Baidu applies router orthogonalization loss to push experts toward distinct specializations rather than converging on copies of each other, and token-balanced loss to prevent any single expert from monopolizing traffic during training. Together they keep the sparse activation stable and genuinely diverse.
How the 128K window was built
The long context was trained in, not bolted on afterward. Baidu progressively scaled the Rotary Position Embedding (RoPE) frequency base from 10K up to 500K across training, with FlashMask attention and memory-efficient scheduling making the longer sequences computationally tractable. Training started at 8K context on text-only data, then expanded to 128K, followed by multi-round reinforcement learning to sharpen the extended reasoning chains this checkpoint produces. Vision stages from the standard ERNIE 4.5 recipe were skipped entirely.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.