Skip to content
RCreddit.com·
Not on the current live radar

85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri

AI summary

A developer has created "Overspill," a disk tier for FreeToken, enabling the execution of large Mixture-of-Experts (MoE) models like the 85 GB DeepSeek-V4-Flash on hardware with limited RAM, specifically a 12 GB RTX 3060 and 64 GB DDR5 RAM. This setup achieves approximately 3 tokens/second. The project was inspired by Colibri and aims to address the challenge of running models that exceed available memory, with performance comparisons against llama.cpp and Colibri also provided.

Why this one

This report details a novel disk-tiering approach, Overspill, allowing an 85 GB model to run on a 12 GB GPU, unlike typical in-memory execution.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 27, 2026, 10:00 UTC

Ingested
Sep 27, 2026, 10:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com