Simulating fault tolerance with stage skipping in pipeline-parallel training [R]
Templar's recent work explores fault tolerance in Crucible, their distributed pre-training platform, aiming to maintain training with healthy workers when a pipeline stage fails. Simulations with a 178M model, eight replicas, and four stages per replica showed that a 1% per-replica failure probability per global step, where each outage removed a stage for six global steps, resulted in validation loss staying close to the no-failure baseline. Each configuration was compared against its own no-failure run.
This report uniquely details how Templar's simulations, unlike prior work, specifically test fault tolerance by skipping stages with fixed projections.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 22, 2026, 23:01 UTC
- Ingested
- Sep 22, 2026, 23:01
- Source type
- Dev community
Full text isn't available here.
Read at source →