Ai2 releases Olmo-core 3, an open MoE training stack benchmarked at 1.2T parameters
Ai2 released Olmo-core 3, an open training framework for large mixture-of-experts models. The company reports about 2.7× the throughput of its earlier FSDP-based implementation and benchmarks up to 1.2 trillion total parameters on 512 NVIDIA B300 GPUs. It is the infrastructure for Ai2's next MoE-based Olmo, not a model release.
Ai2 has released Olmo-core 3, an open training framework for mixture-of-experts (MoE) models. According to Ai2, it scales into the trillion-parameter range. The release includes a tech report, code and an interactive demo. This is training infrastructure, not a new model. Ai2 says it underpins the next generation of Olmo, which will use an MoE architecture.
What changed
Olmo-core's earlier MoE implementation used fully sharded data parallelism (FSDP), which gathers and reshards weights for each small batch. Olmo-core 3 switches to a distributed data parallelism (DDP)-based design. Experts stay resident on GPUs and the relevant data is routed to them.
The stack combines three parallelism techniques:
- Expert parallelism: spreads experts across GPUs.
- Pipeline parallelism: splits model layers across GPU groups.
- Distributed optimizer: shards optimizer state instead of replicating it on every GPU.
Three routing and compute optimizations sit on top:
- Rowwise expert parallelism: places routed data directly into expert input buffers.
- GPU-resident routing: keeps routing metadata on the GPU so the CPU does not wait on copies.
- Grouped GEMM: combines many small expert computations.
The framework also supports MXFP8, a lower-precision number format.
Reported numbers
All figures below are Ai2's own measurements, and several are explicitly preliminary or system-only tests.
- Expert scaling: Ai2 raised the expert pool from 8 to 128 with four experts selected per token, keeping active parameters near 3.2B. Total parameters grew from 4.6B to 47B, and training throughput fell by less than 5%.
- Versus earlier FSDP stack: In a preliminary test on eight NVIDIA B300 GPUs, a 47B-parameter MoE processed 52,000 tokens per second per GPU, against 19,400 for the earlier implementation, about 2.7×.
- MXFP8: On four B300 GPUs with uniform expert load, MXFP8 gave about 21% higher throughput than a BF16 baseline. Peak active memory fell from 103 GiB to 95 GiB. Ai2 says most of the gain came from feed-forward computation and inter-expert data movement, not attention.
- Trillion-scale: A 1.2-trillion-parameter model with 58.36B active parameters per token ran across 512 B300 GPUs. Peak observed throughput was 858 TFLOP/s/GPU.
- DeepEP v2: An experimental configuration reached 2.38 trillion total parameters. Ai2 describes this as a short-capacity test, not sustained training.
The trillion-scale tests used random routing to measure system performance, so they say nothing about the quality of a trained model.
Findings in the tech report
The report also documents negative and unexpected results:
- A load-balancing score can improve while actual workload balance worsens. Ai2 calls this "token gerrymandering."
- Lowering expert learning rates to compensate for fewer tokens per expert did not improve results in the model family tested.
- GPU kernel timing varied with input values even when matrix dimensions were identical, so benchmarks need matched values as well as matched shapes.
- Overlapping communication and computation on separate GPU streams sometimes slowed end-to-end execution.
Context
OlmoE used 64 routed experts. Olmo 3 was dense, and its training stack was built around that design. Ai2 positions Olmo-core 3 against NVIDIA's Megatron-Core, an established option for large MoE training, as an integrated stack within the Olmo framework. It does not publish a head-to-head comparison with Megatron-Core. Ai2 says the next Olmo will use an MoE architecture, a larger dataset and its longest context window yet. Parameter count, context length, release date and pricing for that model are not yet disclosed.
What this means
The notable contribution is the open, documented engineering around MoE training, not a headline benchmark. The sub-5% throughput loss when total parameters grow roughly 10× at fixed active compute is the key claim. If it holds in real training runs, it makes high-expert-count sparsity cheaper for groups without proprietary stacks. The negative results, especially on load-balancing metrics and communication overlap, are useful to anyone building similar systems.
Caveats apply. The headline scale figures use random routing and a short capacity test, and the 2.7× gain comes from a preliminary eight-GPU run. Real trillion-parameter training quality and stability remain unproven until Ai2 ships the model built on this stack.
Related Articles
Anthropic Red Team: GLM-5.3 Matches Claude on Binary Exploitation for First Time
Anthropic's Frontier Red Team reports that Zhipu AI's GLM-5.3 achieved full control flow hijacks in 4% of binary exploitation trials, versus 6% for Claude Mythos Preview. Predecessor models Claude Opus 4.6 and GLM-5.2 scored zero, marking what Anthropic calls a crossed threshold in offensive cyber capability.
Stanford, Caltech Researchers Wire GPT-6 Astra Directly Into a Robot to Clean an Unfamiliar Kitchen
Researchers built HomeBody, a system that connects GPT-6 Astra directly to a Unitree G1 robot's skill library, letting it explore, map, and tidy an unfamiliar kitchen without a trained control layer in between. The team reports latency, overheating servos, and compute cost as current limitations.
Study: Access to AI Advice Nearly Eliminates People's Willingness to Say "I Don't Know"
A five-study research project with 3,132 participants found that access to AI advice—even from a model that was mostly wrong—nearly wiped out people's willingness to admit uncertainty. Confidence rose sharply while accuracy fell.
Nvidia's SoL-Pi Cuts Coding Agent Token Usage by Up to 49% Through Automated Harness Optimization
A new Nvidia research system called SoL-Pi automatically rewrites the control logic of coding agents rather than the underlying model, cutting token usage by up to 49% while keeping performance nearly intact. The approach could shift efficiency gains in AI agents from model-level tricks to harness-level engineering.
Comments
Loading...