ID Balancing Cuts Routing Overload by 89% in 70B Sparse AI Models

A control-theory take on expert routing cuts load imbalance in sparse Mixture-of-Experts training by over 50 percent at extreme sparsity.

·
·
ID Balancing Cuts Routing Overload by 89% in 70B Sparse AI ModelsPRO
  • ID Balancing reframes MoE load balancing as PID control, treating existing methods as incomplete controllers.
  • DeepSeek's loss-free method is an integral controller; Kimi's Quantile Balancing is proportional.
  • Adds magnitude-scaled integral and worsening-gated derivative terms for stronger corrections during spiraling imbalance.
  • Cuts worst-case MaxVio by over 50 percent at Top-3-of-768 routing versus best baselines.
  • Scales from 18.9B to 69.9B parameters with nearly unchanged load imbalance.
  • Maintains competitive language modeling scores; no public code release yet.

Training a highly sparse Mixture-of-Experts model becomes harder as the expert pool grows and the number selected per token shrinks. A small routing skew can send excess traffic to a few experts while leaving others undertrained. A new paper introduces ID Balancing, a control-theory approach that stabilizes routing in configurations reaching 768 experts and 69.9 billion parameters.

Why extreme sparsity destabilizes routing

Mixture-of-Experts (MoE) layers increase parameter count without a proportional increase in per-token computation. A learned router scores every expert, then sends each token to only the top K. In a Top-3-of-768 configuration, for example, each token activates 0.4 percent of the available experts.

That efficiency depends on distributing tokens evenly. Popular experts can exceed their processing capacity and drop assignments, while lightly used experts receive too little training. In expert-parallel systems, where experts run on separate accelerators, overloaded workers delay the entire step as other devices sit idle. Persistent concentration can develop into routing collapse.

Auxiliary load-balancing losses address the problem by penalizing uneven assignments, but their gradients modify the language-modeling objective. DeepSeek-V3 instead adjusts a per-expert router bias outside that objective: overloaded experts receive a lower bias, and underused experts receive a higher one. Kimi’s Quantile Balancing uses a related bias update. The paper finds that both approaches lose effectiveness as routing becomes more sparse.

Router bias becomes a feedback loop

The paper models bias-based routing as a feedback controller. The system measures each expert’s load, compares it with the target average, updates the routing bias, and observes the resulting distribution during later steps. The bias therefore acts as controller state rather than a learned parameter updated through backpropagation.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads