RCreddit.com·
Not on the current live radar
Lessons learned while building Apex-2
The developer of Apex-2 shared lessons learned from building the model, noting that each GPU achieved approximately 40% MFU. Training on two GH200 instances, which cost $2.29/hour per GPU, ran about 1.9x faster than on a single GPU at a lower price, despite H100 instances costing $4.19/hour per GPU. Merging every 350 steps, with merges typically taking less than a minute, contributed to this efficiency.
This report offers a rare, specific comparison of training costs and speeds between H100 and GH200 GPUs, unlike typical high-level discussions.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 5, 2026, 11:00 UTC
- Ingested
- Oct 5, 2026, 11:00
- Source type
- Dev community
Full text isn't available here.
Read at source →