Notableinfrastructure computeHugging Face

Olmo-core 3 Launches with 47B Parameter Capacity and 2.7× Throughput Improvement

Published
Oct 1, 2026 — 15:01 UTC

Olmo-core 3, an open mixture-of-experts (MoE) training system, increases parameter capacity from 4.6 billion to 47 billion. This new system processes approximately 52,000 tokens per second per GPU, a significant rise from the 19,400 tokens achieved with the previous implementation. The training throughput has decreased by less than 5%, while the overall throughput improvement factor is about 2.7×. Notably, the total parameters in the benchmark reach 1.2 trillion, with active parameters per token at 58.36 billion. The highest observed throughput is 858 TFLOP/s per GPU, and the total parameters in the DeepEP v2 test amount to 2.38 trillion. Additionally, training throughput has increased by about 21% with the MXFP8 format, and peak active memory has been reduced from 103 GiB to 95 GiB. Olmo-core 3 employs distributed data parallelism (DDP) as opposed to the fully sharded data parallelism (FSDP) used in earlier versions, allowing it to scale MoE training into the trillion-parameter range while maintaining computational efficiency. This follows previous advancements in MoE training systems, reflecting ongoing improvements in large-scale AI model training.

Summarised from Hugging Face Blog's original report by the Turing Wire Newsdesk. Read the original for the full story.

Source: Hugging Face Blog