HChuggingface.co·
Not on the current live radar
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Olmo-core 3 is an open and scalable training infrastructure designed for large Mixture-of-Experts (MoEs) models. Benchmarked on NVIDIA B300 GPUs, it achieved a throughput of 858 TFLOP/s/GPU with a 1.2-trillion-parameter model, using 58.36 billion parameters active per token across 512 GPUs. These benchmarks focused on system performance using random routing, not the quality of a trained model. Further details on its design and experiments are available in its technical report and on GitHub.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 1, 2026, 16:00 UTC
- Ingested
- Oct 1, 2026, 16:00
- Source type
- Official
Full text isn't available here.
Read at source →