跳到正文
RCreddit.com·
暂不在当前实时榜单

Running Qwen3.8-Flash-Next locally on a 12GB VRAM card

AI 摘要

A user successfully ran Qwen3.8-Flash-Next on an RTX 4070 12GB VRAM card, achieving over 20 tokens/second. This was accomplished by applying PR #28243, which enables a 1.78 GB shared-Q4_K_M compact head, and using the -ncmoe 45 setting. This configuration resulted in 77–96% acceptance rates across various tasks like coding, summarization, and creative generation.

为什么是这条

This report uniquely details the specific PR (#28243) and configuration (-ncmoe 45) that enabled Qwen3.8-Flash-Next to break the 20 t/s barrier on 12GB VRAM, unlike general performance claims.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月15日 01:01 UTC

收录
2026年9月15日 01:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com