Will Nvidia Vera Rubin actually make LLM pre-training faster? And are 10T+ parameter models next?
A discussion on Reddit questions whether Nvidia's Vera Rubin platform will significantly accelerate LLM pre-training, noting that many advertised gains appear linked to low-precision formats and inference rather than pre-training. The conversation also explores the potential for 10T+ parameter models, considering if data, power, and cost limitations will shift focus towards Mixture-of-Experts (MoE) and improved data quality over simply increasing model size.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月29日 14:34 UTC
收录当时偏移:UTC+02026年9月30日 01:00 UTC
- 发布
- 2026年9月29日 14:34
- 收录
- 2026年9月30日 01:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
Nvidia's Vera Rubin numbers look huge, but most of the headline gains seem tied to low-precision formats and inference. How much of that actually carries over to pre-training?
Also curious whether we'll start seeing 10T+ parameter models, or if data, power, and cost are the real limits now, and the focus stays on MoE and better data instead of just going bigger.
Anyone with hardware or training experience, I'd love to hear your take.