Ai2 Rebuilds OLMo-Core 3 to Train Trillion-Parameter AI 2.7x Faster
Ai2 open-sources Olmo-core 3, a redesigned MoE training stack that hits 2.7x throughput and scales past a trillion parameters.
- Ai2 released Olmo-core 3, open MoE training infrastructure scaling to trillion-parameter models.
- Delivers 2.7x throughput over Olmo-core 2 on a 47B MoE across 8 B300 GPUs.
- Switches from FSDP to DDP, keeping experts resident on GPUs instead of regathering weights.
- Benchmarked at 1.2T parameters on 512 GPUs hitting 858 TFLOP/s/GPU with random routing.
- MXFP8 support lifts throughput ~21% versus BF16 while cutting peak memory from 103 to 95 GiB.
- Tech report documents failures like token gerrymandering and misleading load-balance scores.
Ai2 rebuilds Olmo’s training stack for trillion-parameter MoEs
Ai2, the Allen Institute for AI, has released Olmo-core 3, a fully open training stack for large mixture-of-experts models. The organization plans to use it for the next generation of Olmo models.
Mixture-of-experts models contain many feed-forward expert networks, while a router sends each token to only a small selection. This sparse activation lets parameter capacity grow without a proportional increase in computation. At cluster scale, token exchange, weight movement, optimizer state, and routing coordination can consume those savings. Olmo-core 3 is designed to keep throughput nearly flat as the expert pool expands.
A 2.7× gain on eight B300s
Ai2 reports two preliminary benchmarks that show how the stack handles both implementation overhead and rising expert counts.
| Test | Configuration | Result |
|---|---|---|
| Training stack | 47B-parameter MoE on eight NVIDIA B300 GPUs | 52,000 tokens per second per GPU, up from 19,400 with the earlier FSDP implementation, or about 2.7× higher throughput |
| Expert scaling | Expert pool increased from 8 to 128, with four experts selected per token and about 3.2B active parameters | Total capacity increased from 4.6B to 47B parameters while throughput declined by less than 5% |
Resident experts cut weight traffic
Ai2’s earlier MoE implementation used fully sharded data parallelism, or FSDP, configured to gather and reshard expert weights for each small batch. Once experts outnumber GPUs, repeatedly moving those weights across the interconnect can leave compute units waiting for data.
Olmo-core 3 combines a DDP-based design with expert parallelism. Experts remain resident on assigned GPUs, and token data travels to the devices holding the selected experts. This layout avoids gathering and resharding the full set of expert weights for every batch.
The cluster is split three ways
Fitting a trillion-parameter model requires partitioning expert weights, model layers, and optimizer state across the cluster:
- Expert parallelism assigns different experts to different GPUs, so each device stores only part of the expert pool.
- Pipeline parallelism assigns groups of model layers to different GPU stages, reducing the weights held by each device.
- A distributed optimizer shards optimizer state across GPUs, avoiding a full copy on every device.
Routing work stays on the GPU
Communication speed also depends on how routed tokens are packed, transferred, and presented to each expert. Olmo-core 3 adds three optimizations for that path:
- Rowwise expert parallelism writes routed token rows directly into expert input buffers, reducing rearrangement and copying.
- GPU-resident routing keeps routing metadata on the device, allowing the CPU to queue work without waiting for a copy back from the GPU.
- Grouped GEMM combines many small expert matrix multiplications into grouped operations that use the GPU more efficiently.
MXFP8 reduces compute and memory costs
Olmo-core 3 supports MXFP8, an 8-bit floating-point format that applies separate scaling to small blocks of values. On four NVIDIA B300 GPUs with work distributed uniformly across experts, selective MXFP8 use increased training throughput by about 21% over a BF16 baseline. Peak active memory fell from 103 GiB to 95 GiB.
Feed-forward computation and expert communication supplied most of the improvement, with a smaller contribution from attention. The measurement used uniform routing; learned routers can produce uneven workloads. Full-run convergence and final model quality remain separate validation questions.
A 1.2-trillion-parameter systems test
Ai2 also tested a model with 1.2 trillion total parameters and 58.36 billion active parameters per token across 512 GPUs. The run reached a peak of 858 TFLOP/s/GPU, counting useful model computation. Random routing isolated system performance, leaving training quality outside the scope of the test.
Using DeepEP v2 for expert communication, the team also reached 2.38 trillion total parameters. That figure comes from a short capacity test, with sustained-training behavior still unreported.
Four failure modes from the report
The technical report documents several negative results that affect implementation and benchmarking:
- A routing-balance score improved even as the actual workload became less balanced, a failure mode the team calls token gerrymandering.
- Reducing expert learning rates to account for the smaller number of tokens processed by each expert did not improve results in the tested model family.
- GPU execution times changed with input values even when matrix dimensions stayed fixed, so reliable performance comparisons require matching inputs as well as matching shapes.
- Overlapping communication and computation on separate GPU streams sometimes increased end-to-end training time.
Large expert pools are the best fit
Olmo-core 3 is most relevant when teams need to expand expert counts while keeping active parameters stable, combine pipeline and expert parallelism to fit a model, or inspect and modify the complete training implementation. Dense models and small MoEs that already fit comfortably on a cluster offer less reason to migrate from Megatron-Core or an established FSDP stack.
The source code is available on GitHub. Ai2 has also published an interactive walkthrough showing how data, expert, and pipeline parallelism compose from one GPU to 512 GPUs.
Ai2 is treating training infrastructure as part of its open-model release strategy. Publishing the stack gives researchers access to the implementation choices, benchmarks, and failed experiments behind its next Olmo generation. The release removes a software barrier for academic groups and smaller labs, while the hardware demands of the largest configurations remain substantial.