跳到正文
RCreddit.com·
暂不在当前实时榜单

Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM

AI 摘要

An inference engine called InferredThoughts has been developed for MoE models that exceed VRAM and RAM capacity. This engine allows models like Qwen3.8-Flash-Next 177B NVFP4(119GiB) to run on a single 16 GB RTX 5060 Ti with 32 GB RAM, achieving 9-10 tok/s via SSD streaming. Most of the model resides on the SSD, with experts loaded on demand. This setup utilizes 20 GiB in memory and 99 GiB on SSD, including 48.5 GiB for routed experts and a 50.7 GiB n-gram table.

为什么是这条

This report details a novel inference engine that allows large MoE models to run on consumer-grade GPUs by streaming most of the model from SSD, unlike traditional methods requiring full in-memory loading.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月28日 00:00 UTC

收录
2026年9月28日 00:00
来源类型
开发者社区

讨论趋势

→ 平稳
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

本站未收录正文。

前往源站阅读 →
来源·reddit.com