Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

AI summary

A developer reported significant performance improvements for the Qwen3.8-Flash-Next model on a 12GB RTX 5070. Initially achieving 15 tok/s output and 100-120 tok/s prompt processing with IQ3_XXS quant using llama.cpp, they later developed a custom inference engine. This engine boosted performance to ~65 tok/s output and ~430 tok/s prompt processing with the same IQ3_XXS quant, with 2-bit quants using RCO-GSQ quantization running even faster.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 24, 2026, 20:01 UTC

Ingested
Sep 24, 2026, 20:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com