Olmo-core 3 released: open MoE training stack benchmarked past 1 trillion parameters

The Olmo team has released Olmo-core 3, a significant upgrade to its framework for developing large language models, built around a redesigned open mixture-of-experts (MoE) training system. The stated aim is to scale MoE training into the trillion-parameter range while preserving computational efficiency. The team calls it one of the core systems behind the next generation of Olmo.
The problem it targets: MoE models can hold many more parameters without every input using all of them, but the full model still has to be stored across GPU memory and updated during training, and routing inputs to the right experts across a cluster adds communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage. In one benchmark, the team raised the expert pool from 8 to 128 while still selecting four experts per token, keeping active parameters per token at about 3.2B. Total parameters grew from 4.6B to 47B, while training throughput fell by less than 5%.
The main design change is in parallelism. Olmo's earlier MoE implementation used fully sharded data parallelism (FSDP), configured to gather and reshard weights for each small batch. Olmo-core 3 switches to a system based on distributed data parallelism (DDP), which keeps experts resident on GPUs and routes the relevant data to them. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, against 19,400 with the earlier implementation, about 2.7× the throughput. The post notes that NVIDIA's Megatron-Core is an established option for training large MoEs.
Three techniques split the model and its training state across hardware: expert parallelism spreads experts across GPUs; pipeline parallelism splits the model's layers across groups of GPUs; and a distributed optimizer spreads optimizer state instead of keeping a full copy on every GPU. Further optimizations cut routing and compute costs: rowwise expert parallelism places routed data directly into expert input buffers, GPU-resident routing keeps routing metadata on the GPUs so the CPU does not wait for copies, and grouped GEMM combines many small expert computations.
Olmo-core 3 also supports MXFP8, a lower-precision number format. In a controlled benchmark on four B300 GPUs with work spread uniformly across experts, enabling MXFP8 where it helped most raised end-to-end training throughput by about 21% over a BF16 baseline, and peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and data movement between experts rather than attention alone.
At scale, the team benchmarked configurations on B300 GPUs including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. Its highest observed throughput was 858 TFLOP/s/GPU. These tests used random routing, so they measure system performance, not the quality of a trained model. Separately, experiments with DeepEP v2, an alternative way of handling communication between experts, reached a configuration with 2.38 trillion total parameters. That was a short-capacity test rather than a full training run, so it shows reachable scale, not sustained training performance.
The accompanying technical report documents experiments and findings, including approaches the team tested and chose not to adopt. A score meant to encourage balanced routing could improve even as the real workload became less balanced; the team calls this failure token gerrymandering. Lowering experts' learning rates because they see fewer tokens did not improve results in the model family tested. GPU calculations took different amounts of time when the input values changed, even with identical matrix dimensions, so performance comparisons need matching values as well as matching shapes. And overlapping communication and computation on separate GPU streams did not always help; in some tests it slowed end-to-end execution.
The team says its next-generation Olmo will use an MoE architecture, and it is aiming for its most capable Olmo yet, trained on its largest dataset with its longest context window. Olmo 3 used a dense architecture, and the earlier OlmoE used 64 routed experts. Olmo-core 3 is described as fully open: researchers and developers can use it to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism and other parts of the system.
Key facts
- Olmo-core 3 moves MoE training from the earlier FSDP-based implementation to a DDP-based design that keeps experts resident on GPUs; in a preliminary test on eight B300 GPUs a 47B MoE reached 52,000 tokens/s/GPU versus 19,400, about 2.7×.
- In one benchmark the expert pool grew from 8 to 128 (four experts per token, about 3.2B active parameters), total parameters rose from 4.6B to 47B, and throughput fell by less than 5%.
- MXFP8 gave about 21% higher end-to-end throughput than BF16 on four B300 GPUs and cut peak active memory from 103 GiB to 95 GiB.
- The largest benchmark was a 1.2-trillion-parameter model (58.36B active per token) on 512 B300 GPUs, peaking at 858 TFLOP/s/GPU with random routing; a short DeepEP v2 test reached 2.38 trillion total parameters.
- The stack is fully open, and the next-generation Olmo will use an MoE architecture.
Why it matters
Training large models costs a lot of compute and energy, and the post says that puts advanced model development out of reach for many academic researchers and smaller labs. MoE models promise more capacity per unit of compute, but communication and coordination costs can eat that advantage as they grow. Olmo-core 3 is built to close that gap, and the near-flat throughput when total parameters grew from 4.6B to 47B is the clearest evidence offered. The team also argues that model weights are more useful when the infrastructure and training decisions behind them are open too.
Who it affects
The post names researchers and developers who want to train their own MoEs, including academic researchers and smaller labs. It also affects anyone following the next generation of Olmo, which will be an MoE built on this stack. Teams currently using NVIDIA's Megatron-Core, described in the post as an established option for large MoEs, now have another integrated stack to look at.
How to use it
The post says researchers and developers can use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism and other parts of the system. It points to the technical report for the systems design, experiments and ablations, to the code on GitHub, and to an interactive walkthrough of how data, expert and pipeline parallelism work together from a single GPU to many. The benchmarks were run on NVIDIA B300 GPUs. No licence terms or GitHub repository URL are given in the text.
How solid is it
The numbers come from the team's own benchmarks, and the post is careful about their scope. The 2.7× figure is from a preliminary test on eight B300 GPUs. The MXFP8 gain was measured on four B300 GPUs with work spread uniformly across experts. The 1.2-trillion-parameter and 858 TFLOP/s/GPU results used random routing, and 858 is the highest observed throughput, not a typical one. The 2.38-trillion-parameter configuration was a short-capacity test rather than a full training run. No throughput comparison against NVIDIA Megatron-Core is given.
Risks and caveats
No trained model of trillion-parameter size is reported, so these results say nothing about model quality at that scale. The post does not say how many GPUs were used in the 4.6B to 47B expert-pool benchmark. No cost, energy or training-time figures are given for the benchmarks. The team's own findings also warn that optimizations interact: overlapping communication and computation sometimes slowed training, and a balance score could improve while the real workload became less balanced. No release date, size or benchmark results are given for the next-generation Olmo.
“model weights are more useful when the infrastructure and training decisions behind them are open too.”
— Olmo-core 3 announcement post