跳到正文
RCreddit.com·
暂不在当前实时榜单

Lessons learned while building Apex-2

AI 摘要

The developer of Apex-2 shared lessons learned from building the model, noting that each GPU achieved approximately 40% MFU. Training on two GH200 instances, which cost $2.29/hour per GPU, ran about 1.9x faster than on a single GPU at a lower price, despite H100 instances costing $4.19/hour per GPU. Merging every 350 steps, with merges typically taking less than a minute, contributed to this efficiency.

为什么是这条

This report offers a rare, specific comparison of training costs and speeds between H100 and GH200 GPUs, unlike typical high-level discussions.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月5日 11:00 UTC

收录
2026年10月5日 11:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com