Skip to content
HChuggingface.co·
Not on the current live radar

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

AI summary

Olmo-core 3 is an open and scalable training infrastructure designed for large Mixture-of-Experts (MoEs) models. Benchmarked on NVIDIA B300 GPUs, it achieved a throughput of 858 TFLOP/s/GPU with a 1.2-trillion-parameter model, using 58.36 billion parameters active per token across 512 GPUs. These benchmarks focused on system performance using random routing, not the quality of a trained model. Further details on its design and experiments are available in its technical report and on GitHub.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 1, 2026, 16:00 UTC

Ingested
Oct 1, 2026, 16:00
Source type
Official

Full text isn't available here.

Read at source →