RCreddit.com·
暂不在当前实时榜单
85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri
A developer has created "Overspill," a disk tier for FreeToken, enabling the execution of large Mixture-of-Experts (MoE) models like the 85 GB DeepSeek-V4-Flash on hardware with limited RAM, specifically a 12 GB RTX 3060 and 64 GB DDR5 RAM. This setup achieves approximately 3 tokens/second. The project was inspired by Colibri and aims to address the challenge of running models that exceed available memory, with performance comparisons against llama.cpp and Colibri also provided.
This report details a novel disk-tiering approach, Overspill, allowing an 85 GB model to run on a 12 GB GPU, unlike typical in-memory execution.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月27日 10:00 UTC
- 收录
- 2026年9月27日 10:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →