85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri
A developer has created "Overspill," a disk tier for FreeToken, enabling the execution of large Mixture-of-Experts (MoE) models like the 85 GB DeepSeek-V4-Flash on hardware with limited RAM, specifically a 12 GB RTX 3060 and 64 GB DDR5 RAM. This setup achieves approximately 3 tokens/second. The project was inspired by Colibri and aims to address the challenge of running models that exceed available memory, with performance comparisons against llama.cpp and Colibri also provided.
This report details a novel disk-tiering approach, Overspill, allowing an 85 GB model to run on a 12 GB GPU, unlike typical in-memory execution.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 27, 2026, 10:00 UTC
- Ingested
- Sep 27, 2026, 10:00
- Source type
- Dev community
Full text isn't available here.
Read at source →