OZZZER · AI NEWS1 of 3 free stories opened
← Back to AI News

Unclassified · 1 Oct 2026 · 17:01 CEST

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Hugging Face · 1 Oct 2026 · 17:01 CESTRead original at Hugging Face ↗
Share
LinkedInXFacebookWhatsApp
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Publisher preview · OZZZER analysis pending editorial review.

Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.

Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach—they can contain many more learned components, or parameters, without requiring every input to use all of them.

But the full model still has to be stored across GPU memory and updated during training, and directing inputs to the right experts – the specialized components within an MoE – across a cluster creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only…

Excerpt supplied by the publisher.

Source

Hugging Face · 1 Oct 2026 · 17:01 CEST

Open the original at Hugging Face ↗