Skip to content
RCreddit.com·
Not on the current live radar

Running Qwen3.8-Flash-Next locally on a 12GB VRAM card

AI summary

A user successfully ran Qwen3.8-Flash-Next on an RTX 4070 12GB VRAM card, achieving over 20 tokens/second. This was accomplished by applying PR #28243, which enables a 1.78 GB shared-Q4_K_M compact head, and using the -ncmoe 45 setting. This configuration resulted in 77–96% acceptance rates across various tasks like coding, summarization, and creative generation.

Why this one

This report uniquely details the specific PR (#28243) and configuration (-ncmoe 45) that enabled Qwen3.8-Flash-Next to break the 20 t/s barrier on 12GB VRAM, unlike general performance claims.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 15, 2026, 01:01 UTC

Ingested
Sep 15, 2026, 01:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com