The release of Olmo-core 3 represents a significant upgrade to the framework for developing large language models, featuring a redesigned open mixture-of-experts (MoE) training system. The framework is engineered to scale MoE training into the trillion-parameter range while preserving computational efficiency.

Shifting from Dense to Sparse Infrastructure

While Olmo 3 utilized a dense architecture, Olmo-core 3 extends the framework for much larger MoE models. The system has transitioned from fully sharded data parallelism (FSDP) to a design based on distributed data parallelism (DDP). This architecture keeps experts resident on GPUs and routes relevant data to them, avoiding the repeated weight gathering and resharding characteristic of previous implementations.

Performance Benchmarks and Optimization

In preliminary tests on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU—approximately 2.7 times the throughput of the earlier FSDP-based stack. The infrastructure has been benchmarked at scales reaching 1.2 trillion total parameters, achieving a throughput of 858 TFLOP/s/GPU.

Key optimization techniques integrated into Olmo-core 3 include:

  • Expert Parallelism: Spreading experts across the GPU cluster.
  • Pipeline Parallelism: Splitting model layers across GPU groups to manage memory.
  • MXFP8 Support: Utilizing lower-precision number formats to reduce data movement and computation costs.

Technical Insights and Challenges

The associated technical report documents several critical findings from the development process. Researchers identified a phenomenon called "token gerrymandering," where routing balance scores could improve even as actual workloads became less balanced. Furthermore, the team noted that overlapping communication and computation on separate GPU streams did not always increase end-to-end training speed, highlighting the complex trade-offs in high-performance AI training.

Olmo-core 3 serves as the foundation for the next generation of Olmo models, which are expected to feature larger datasets and longer context windows.