GitHub Repo
Ai2 Open-Sources Olmo-core 3, Redesigning Expert Placement and Data Routing for MoE Training
The new version keeps experts resident on GPUs and uses data routing to reduce the cost of repeatedly aggregating weights. Initial tests reported about 2.7× throughput. A trillion-parameter configuration has passed systems testing, but it used random routing, so training quality and stability remain unverified.

Ai2 released Olmo-core 3 on October 1, redesigning the mixture-of-experts (MoE) training system in its open-source framework. The update focuses on expert placement and cross-GPU data movement, and will serve as the training foundation for the next generation of Olmo. The code is publicly available for researchers to modify and experiment with. Official announcement
The key change is a shift from an implementation based primarily on fully sharded data parallelism (FSDP) to one based on distributed data parallelism (DDP): expert weights remain on the GPUs, while token data is sent to the corresponding experts. This reduces the cost of repeatedly aggregating and resharding weights for each small batch. In tests on eight NVIDIA B300 GPUs, the team measured a 47-billion-parameter MoE model. Throughput per GPU rose from 19,400 tokens per second to 52,000, about 2.7× that of the previous implementation. These are preliminary results and should not be assumed to apply directly to other clusters. Test and architecture details
The limits of its scalability claims are also clear. The team tested a configuration with 1.2 trillion total parameters and 58.36 billion parameters active per token on 512 B300 GPUs, using random routing to measure system performance. This result shows that the framework can support large configurations, but does not establish the convergence quality, long-term stability, or actual training cost of a model at that scale. Limitations of the official tests
Putting the system into production still requires handling dependency and environment differences. The repository provides the ai2-olmo-core package and instructions for installing from source; some features require additional kernel packages. The README specifically notes that grouped_gemm, used by dropless MoE, may need to be compiled manually. The official Docker image includes core and optional dependencies, but does not come with Olmo-core itself preinstalled. Different hardware, drivers, or CUDA versions may also require rebuilding the image. Repository installation instructions
For teams considering the new version, the next step is to compare throughput, peak memory usage, and training loss on their own hardware and with realistic routing workloads, then decide whether the architectural changes are worth adopting. This is an engineering recommendation based on the scope of the tests: publicly available code lowers the barrier to inspection and modification, but benefits across environments still need to be reproduced.