Training a 3.8B LLM to 0.384 CORE for $998
A researcher successfully trained a 3.8B LLM to 0.384 CORE for $998, demonstrating that meaningful models can be trained by individuals with limited budgets. The training involved 25,000 steps and 57.3B tokens, taking 35.9 hours. The process achieved a steady state of approximately 480,000 tokens/sec. Initial attempts at document-boundary masking with flex attention were discarded, as cross-document leakage was found not to significantly worsen performance, leading to a simpler best-fit packing approach.
Why this oneThis report uniquely details the specific cost ($998) and performance (0.384 CORE) for training a 3.8B LLM, unlike general discussions of budget-friendly model training.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 10, 2026, 04:00 UTC
- Ingested
- Sep 10, 2026, 04:00
- Source type
- Unclassified
Full text isn't available here.
Read at source →