Ai2 has released Olmo-core 3, an open training framework designed to scale mixture-of-experts models into the trillion-parameter range. These models contain many specialized components but activate only a subset for each token, reducing computation while creating difficult memory and network-coordination problems.
The redesigned stack keeps experts resident on GPUs and routes data to the relevant experts, replacing an earlier approach that repeatedly gathered and redistributed model weights. In an eight-GPU B300 test, Ai2 says a 47-billion-parameter model processed 52,000 tokens per second per GPU, about 2.7 times the throughput of its previous implementation.
Olmo-core 3 combines expert, pipeline and optimizer parallelism with GPU-resident routing and grouped matrix operations. It also supports the lower-precision MXFP8 format. In one four-GPU benchmark, MXFP8 increased end-to-end throughput by about 21 percent and reduced peak active memory from 103 GiB to 95 GiB.
The team tested a 1.2-trillion-parameter configuration across 512 GPUs and briefly reached 2.38 trillion parameters with an alternative communications system. Those runs used random routing or short capacity tests, so they demonstrate infrastructure scale rather than model quality or stable full training. The code is open for researchers to inspect and adapt.