RCreddit.com·
暂不在当前实时榜单
Lessons learned while building Apex-2
The developer of Apex-2 shared lessons learned from building the model, noting that each GPU achieved approximately 40% MFU. Training on two GH200 instances, which cost $2.29/hour per GPU, ran about 1.9x faster than on a single GPU at a lower price, despite H100 instances costing $4.19/hour per GPU. Merging every 350 steps, with merges typically taking less than a minute, contributed to this efficiency.
This report offers a rare, specific comparison of training costs and speeds between H100 and GH200 GPUs, unlike typical high-level discussions.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月5日 11:00 UTC
- 收录
- 2026年10月5日 11:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →