Skip to content
RCreddit.com·
Not on the current live radar

Lessons learned while building Apex-2

AI summary

The developer of Apex-2 shared lessons learned from building the model, noting that each GPU achieved approximately 40% MFU. Training on two GH200 instances, which cost $2.29/hour per GPU, ran about 1.9x faster than on a single GPU at a lower price, despite H100 instances costing $4.19/hour per GPU. Merging every 350 steps, with merges typically taking less than a minute, contributed to this efficiency.

Why this one

This report offers a rare, specific comparison of training costs and speeds between H100 and GH200 GPUs, unlike typical high-level discussions.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 5, 2026, 11:00 UTC

Ingested
Oct 5, 2026, 11:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com