I trained a 44M parameter quantized LLM from scratch on 45B tokens. It ships in 19.8 MB and runs at ~1,900 tok/s on CPU. [P]
A developer trained a 44M parameter quantized LLM, SHADOW-250M, from scratch on 45B tokens, resulting in a 19.8 MB model that runs at ~1,900 tok/s on CPU. While SHADOW-250M performed lower on standard benchmarks like ARC-Easy (0.307) and PIQA (0.570) compared to Supra-50M-Reasoning, it demonstrated superior performance in generating direct and accurate answers to various questions, including jokes, math problems, and date calculations, where Supra often struggled or provided irrelevant information.
This report uniquely details a quantized LLM that, unlike others, prioritizes direct, accurate answers over benchmark scores, showcasing its practical utility in specific tasks.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 15, 2026, 18:01 UTC
- Ingested
- Sep 15, 2026, 18:01
- Source type
- Dev community
Full text isn't available here.
Read at source →