Ai2 opens up its MoE training stack, tested at over a trillion parameters
Olmo-core 3 is infrastructure, not a model, but it is the foundation for the next Olmo and comes with unusually candid negative results.
The Allen Institute for AI (Ai2) on Thursday released Olmo-core 3, a new version of the open-source framework it uses to train its Olmo language models. The headline change is a rebuilt training system for mixture-of-experts (MoE) models, the sparse architecture in which each token only passes through a handful of specialized sub-networks. Ai2 says it has benchmarked the stack at more than one trillion total parameters, and that the next generation of Olmo will be built on it.
This is infrastructure, not a model. There is nothing new to call from an API today. What it does is make public the kind of training plumbing that frontier labs usually keep in-house.
What changed
MoE models promise more capacity for roughly the same compute per token, but the bookkeeping erodes that advantage as they grow: every expert still has to live somewhere in GPU memory, and routing tokens to the right experts across a cluster costs communication. Olmo-core 3 is Ai2's attempt to keep that overhead under control.
The main architectural shift is away from fully sharded data parallelism (FSDP), which in Ai2's earlier MoE implementation gathered and re-sharded weights for each micro-batch, toward a system based on distributed data parallelism (DDP) that keeps experts resident on their GPUs and moves the data to them instead. On top of that, the stack combines three ways of splitting the work:
- Expert parallelism: each GPU holds only part of the expert pool.
- Pipeline parallelism: the model's layers are split across groups of GPUs.
- A distributed optimizer: optimizer state is spread across GPUs rather than copied to each one.
Ai2 also lists lower-level optimizations: placing routed tokens directly into expert input buffers, keeping routing metadata on the GPU so the CPU does not stall waiting for it, grouped GEMM kernels that batch many small expert computations, and support for the lower-precision MXFP8 number format.
The numbers, and what they do and don't show
All figures below come from Ai2's own benchmarks; no independent reproduction is available yet.
- On eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed about 52,000 tokens per second per GPU on the new stack versus 19,400 on Ai2's previous FSDP-based implementation, roughly 2.7x. Ai2 calls this a preliminary test.
- Growing the expert pool from 8 to 128 while still routing each token to four experts kept active parameters around 3.2B, raised total parameters from 4.6B to 47B, and cut training throughput by less than 5%.
- Enabling MXFP8 where it helped most gave about 21% higher throughput than BF16 on four B300s, with peak active memory falling from 103 GiB to 95 GiB.
- A 1.2-trillion-parameter configuration with 58.36B active parameters ran across 512 GPUs, peaking at 858 TFLOP/s per GPU. A separate short test using DeepEP v2 reached a 2.38-trillion-parameter configuration.
The trillion-parameter figures need careful reading. Ai2 says those tests used random routing to measure system performance, not the quality of a trained model, and that the 2.38T run was a short capacity test rather than sustained training. They show that the plumbing can hold a model that size; they say nothing about whether a model trained on it would be any good. Ai2 also positions the stack alongside NVIDIA's Megatron-Core, which it calls an established option for large MoEs, but does not publish a head-to-head comparison.
The more useful part may be the negative results
The accompanying technical report documents several findings that run against common assumptions, which is the kind of detail open releases are uniquely positioned to share:
- A load-balancing score meant to encourage even routing could improve while the real workload became *less* balanced. Ai2 calls this token gerrymandering.
- Lowering experts' learning rates because they see fewer tokens did not improve results in the model family tested.
- GPU kernels took different amounts of time depending on the values being processed, even with identical matrix shapes, so fair performance comparisons need matching inputs, not just matching dimensions.
- Overlapping communication and computation on separate GPU streams sometimes made end-to-end training slower.
What it means for developers
For most people choosing between models today, nothing changes this week. The relevance is longer-range. Ai2 says its next Olmo will be an MoE, trained on its largest dataset and with its longest context window yet, and aims for it to be its most capable Olmo. No release date was given. If that model ships with weights, data and training code all open, it would give developers a frontier-style sparse model whose full training pipeline can be inspected and reproduced, something closed labs do not offer.
For teams that train or fine-tune their own MoEs, Olmo-core 3 is a usable artifact now: Ai2 says the code is open on GitHub and can be adapted to other hardware and routing schemes. Whether its throughput claims hold on clusters other than Ai2's B300 setup is the open question worth testing before adopting it.
Source: Ai2's announcement.
Why this matters
- It opens up MoE training infrastructure at trillion-parameter scale, a layer frontier labs usually keep private.
- Ai2 says its next Olmo will be an MoE with its longest context window yet, built on this stack.
- The documented negative results, such as token gerrymandering, are useful to anyone training sparse models.
Key takeaways
- Olmo-core 3 is training infrastructure, not a new model; nothing changes in APIs today.
- Ai2 reports about 2.7x throughput over its previous FSDP-based MoE stack on eight B300 GPUs.
- The 1.2T and 2.38T parameter runs were system tests with random routing, not trained models.
- All performance figures are Ai2's own; no independent reproduction is available yet.
- No release date has been given for the next-generation Olmo MoE.
Sources
- Ai2 (Allen Institute for AI), via Hugging Face blogPrimaryIntroducing Olmo-core 3: Open, scalable training infrastructure for large MoEshuggingface.co
- open-source
- mixture-of-experts
- training-infrastructure
- olmo
- mxfp8
- nvidia-b300
- Ai2
- NVIDIA
- Olmo