Zyphra Cuts MoE Training Communication Costs by 2.63x Using Router Patterns
Zyphra researchers cut MoE all-to-all communication by up to 2.63x on AMD MI300X GPUs by exploiting predictable token routing patterns, with no changes to the model.
- Zyphra's new paper cuts MoE all-to-all time 1.16-2.63x and step time up to 1.41x.
- MoE all-to-all can eat 45-60% of training step time on AMD MI300X clusters.
- Routers learn correlated patterns early: 0.8% of expert pairs attract 42% of tokens.
- Correlated placement co-locates hot expert pairs, cutting dispatched rows by up to 58%.
- Token shuffling piggybacks on reduce-scatter to pre-position tokens near next-layer experts.
- Model function, routing decisions, and expert weights are unchanged, work is purely a systems win.
At scale, Mixture-of-Experts training can spend more time moving tokens than running expert computation. A paper from Zyphra reports that router choices become predictable early in pretraining, allowing the runtime to co-locate frequently paired experts and move tokens toward likely destinations.
Benchmarks in the paper, run with Megatron-LM on AMD Instinct MI300X clusters, accelerated all-to-all exchanges by 1.16× to 2.63× and the full training step by up to 1.41×.
MoE’s network bottleneck
In an MoE layer, a router assigns each token to its top-k experts. Expert parallelism, or EP, distributes those experts across GPU ranks; EP32 means the expert group spans 32 GPUs. Each forward pass dispatches tokens to the selected experts and returns the resulting activations, while backpropagation requires corresponding data movement.
On clusters with eight MI300X GPUs per node, these all-to-all collectives consume about 45% of an EP32 training step with top-2 routing and 60% with top-6 routing. Higher values of k create more destinations per token, and EP groups that cross node boundaries encounter slower interconnects.
Existing MoE systems reduce this cost through communication overlap, replicated experts, or modified routing schemes. Zyphra’s design operates in the placement and dispatch layers, using observed routing patterns to reduce the amount of data crossing the network.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.