Skip to content
RCreddit.com·

Will Nvidia Vera Rubin actually make LLM pre-training faster? And are 10T+ parameter models next?

AI summary

A discussion on Reddit questions whether Nvidia's Vera Rubin platform will significantly accelerate LLM pre-training, noting that many advertised gains appear linked to low-precision formats and inference rather than pre-training. The conversation also explores the potential for 10T+ parameter models, considering if data, power, and cost limitations will shift focus towards Mixture-of-Experts (MoE) and improved data quality over simply increasing model size.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 29, 2026, 14:34 UTC

IngestedOffset at this time: UTC+0Sep 30, 2026, 01:00 UTC

Published
Sep 29, 2026, 14:34
Ingested
Sep 30, 2026, 01:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Nvidia's Vera Rubin numbers look huge, but most of the headline gains seem tied to low-precision formats and inference. How much of that actually carries over to pre-training?

Also curious whether we'll start seeing 10T+ parameter models, or if data, power, and cost are the real limits now, and the focus stays on MoE and better data instead of just going bigger.

Anyone with hardware or training experience, I'd love to hear your take.

Source·reddit.com