Optimizer Rankings Flip as Batch Size Grows, Study Warns ML Engineers
A new study finds no principled hyperparameter scaling rule consistently preserves optimizer performance across batch sizes, undermining single-batch benchmarks.
- New arXiv paper shows optimizer rankings change with batch size, undermining single-batch benchmarks.
- No principled scaling rule for Muon transfers hyperparameters consistently across language and vision tasks.
- SOAP leads from 128K to 1M tokens but falls 0.021 nats behind Shampoo at 2M tokens.
- Muon and Shampoo also swap positions between 512K and 2M-token batches in Modded-NanoGPT runs.
- Noisy quadratic analysis shows correct scaling shifts from square-root to linear with curvature-to-noise ratio.
- Theoretical SDE and bound-minimization rules can lose to the trivial no-scaling baseline at large batches.
Optimizer rankings change with batch size, study finds
A study by Xingyu Dang, Kaiyue Wen, and Sadhika Malladi finds that optimizer rankings can reverse as training batch size changes, even when every contender is retuned for each batch. The paper challenges a common benchmark shortcut: tune an optimizer at one batch size, publish the result, and use a scaling rule to transfer its hyperparameters to larger runs.
The authors report two related findings. No tested scaling rule transfers Muon’s tuned hyperparameters reliably across training settings, and the best optimizer for language-model pretraining changes with batch size. Developers choosing an optimizer from a single-batch benchmark may therefore select a weaker option for their production batch.
The tune-once shortcut
Global batch size is the number of training tokens processed before one optimizer update. Larger batches usually produce less noisy gradient estimates, which changes the learning rate and other hyperparameters that work best. Linear scaling multiplies the learning rate by the batch-size ratio, while square-root scaling uses the square root of that ratio to preserve the approximate scale of random update fluctuations.
Scaling rules make small experiments cheaper because researchers can tune at an affordable reference batch and project the configuration to a larger run. Earlier scaling work reported that AdamW reaches a critical batch size, the point where larger batches deliver diminishing training speedups, sooner than Muon. That study tested Muon with batches of up to 100 million tokens per step. Batch-dependent rankings narrow such conclusions to the regimes in which the optimizers were measured.
Five optimizers, several crossovers
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.