Olmo-core 3 gets a rebuilt mixture-of-experts system meant to push training into the trillion-parameter range, in the open framework Ai2 uses for its Olmo language models. Per Ai2, a preliminary test on eight NVIDIA B300 GPUs had a 47-billion-parameter model processing 52,000 tokens per second per GPU. The old implementation managed 19,400, so the new one is about 2.7 times faster.
A mixture-of-experts model stores many specialist sub-networks but sends each token through only a few, so total size can grow without a matching rise in compute per token. The price is logistics: every expert has to live in GPU memory somewhere, and tokens must be routed across a cluster to reach them. The rebuilt system in Olmo-core is meant to keep that overhead from eating the savings.
The framework's earlier MoE code relied on fully sharded data parallelism, which pulls model weights together and splits them again for each small batch. The new version switches to distributed data parallelism, keeps experts resident on their GPUs and sends the relevant tokens to them instead.
The speed comparison is against Ai2's own previous code. NVIDIA's Megatron-Core is the established option for large MoE training, and the Allen Institute for AI makes no head-to-head claim against it. What it offers instead is an integrated stack inside the open Olmo-core repository that already trains Olmo, which outside researchers can adapt to other hardware.
The accompanying technical report adds a width test. Growing the expert pool from 8 to 128, still choosing four per token, took total parameters from 4.6 billion to 47 billion at a cost of under 5% of training throughput. Active parameters stayed near 3.2 billion, which is why compute per token barely moved. Work was spread evenly across experts in that benchmark, a best case.
At the top end, a 1.2-trillion-parameter configuration across 512 B300 GPUs peaked at 858 TFLOP/s per GPU. Those runs used random routing, so they measure how fast the system moves data, not how well a trained model performs. A shorter capacity test with an optional communication backend called DeepEP v2 reached 2.38 trillion total parameters, which Ai2 describes as a demonstration of reach rather than sustained training.
Ai2 says the next Olmo will be a mixture-of-experts model trained on its largest dataset with its longest context window, and it has given no date. The report also records approaches the team tried and dropped. One is a failure Ai2 calls token gerrymandering, where a routing score meant to encourage balance improves while the real workload gets less balanced. Another is that overlapping communication with computation sometimes slowed end-to-end training. For outside researchers adapting the code, that list of dead ends is as useful as the speedups.













