RCreddit.com·
暂不在当前实时榜单
Running Qwen3.8-Flash-Next locally on a 12GB VRAM card
A user successfully ran Qwen3.8-Flash-Next on an RTX 4070 12GB VRAM card, achieving over 20 tokens/second. This was accomplished by applying PR #28243, which enables a 1.78 GB shared-Q4_K_M compact head, and using the -ncmoe 45 setting. This configuration resulted in 77–96% acceptance rates across various tasks like coding, summarization, and creative generation.
This report uniquely details the specific PR (#28243) and configuration (-ncmoe 45) that enabled Qwen3.8-Flash-Next to break the 20 t/s barrier on 12GB VRAM, unlike general performance claims.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月15日 01:01 UTC
- 收录
- 2026年9月15日 01:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →