Qwen3.8-27B UD-IQ4_XS Heretic + MTP on a 16 GB card with 55-68 tok/s (24gb and 12gb versions available too)
A user successfully ran the uncensored Qwen3.8-27B (llmfan46's Heretic build, MTP head preserved) on an RTX 4080. The UD-IQ4_XS quantization, using 16 GB, achieved 55 tok/s on code and 50 tok/s on prose with MTP enabled, significantly faster than the 29 tok/s without MTP. Other versions, including 24 GB and 12 GB, are also available, with varying speeds and quality trade-offs.
This report details a specific quantization of Qwen3.8-27B that achieves 55-68 tok/s on a 16 GB card, unlike other versions that offer different speed-quality trade-offs.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月9日 22:54 UTC
收录当时偏移:UTC+02026年10月10日 07:00 UTC
- 发布
- 2026年10月9日 22:54
- 收录
- 2026年10月10日 07:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 同步延迟
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
I wanted the uncensored Qwen3.8-27B (llmfan46's Heretic build, MTP head preserved) on my RTX 4080 with MTP on (because I've been testing TONS of configs, and realized that would be the one)
And none of the published quants were built for that: the good IQ4_XS doesn't leave room for MTP, and the ones that fit are 3-bit.
I was amazed by that. as 16gb is VERY common, so decided to quantize it to fit, work well and loose as less as possible in quality:
I copied Unsloth's per-tensor UD recipe onto llmfan46's BF16, made 12 / 16 / 24 GB versions, and measured KLD against a Q8_0 of the same weights for every quant I could find.
Quality (mean KLD vs Q8_0, wikitext-2 / llama.cpp source code, lower is better)
Quant GiB Prose Code UD-Q5_K_XL (mine, 24 GB) 19.44 0.0045 0.0037 mradermacher i1-IQ4_XS 14.26 0.0202 0.0149 UD-IQ4_XS (mine, 16 GB) 13.27 0.0268 0.0192 llmfan46 Q3_K_M 13.48 0.0639 0.0456 mradermacher i1-IQ3_M 11.89 0.0649 0.0465 UD-IQ3_XXS (mine, 12 GB) 10.18 0.0904 0.0587 mradermacher i1-Q2_K 10.12 0.1551 0.1017 Speed on the 4080 (llama.cpp b11457, 15K-token prompt, 32K ctx, MTP 2 drafts, mean of 3 fixed seeds): UD-IQ4_XS does 55 tok/s on code and 50 on prose, against 29 / 29 without MTP. i1-IQ3_M is faster (70 / 54) at 2.4× the KLD, I prefer quality here over speed, but your call.
- 2 draft tokens beat 3. On the same seeds at 40K: 68 vs 56 tok/s on code, 55 vs 36 on prose. Fewer rejected drafts.
- 48K vs 40K produced the exact same tokens (identical draft-acceptance counts per seed) and was 22% slower. That's pure VRAM spill into shared memory, nothing else.
- The 16 GB edge is brutal. The same 40K config gave 68, 51 and 42 tok/s on code depending only on whether the desktop was holding 1.2, 1.4 or 1.9 GB of VRAM. Opening a Chrome window mid-generation took decode to 0.2 tok/s. If you're on Windows with a monitor on the same card, watch "shared GPU memory" in Task Manager.
- MTP head precision barely matters. q6_K head vs IQ3_S head: 83% vs 83% acceptance on code, 58% vs 54% on prose.
Repo with all three files, the commands, and the scripts (recipe extraction, dry-run verification, KLD, seeded speed bench): https://huggingface.co/codavidgarcia/Qwen3.8-27B-Uncensored-Heretic-MTP-UD-GGUF
The 12 GB and 24 GB context numbers are computed from llama.cpp's reported buffers, not measured on those cards. If you run them, let me know what you get!
Credit to llmfan46 for the weights, Unsloth for the recipes and mradermacher for the imatrix
edit: added the KLD chart since a few people asked about the numbers, (lower is better)
https://preview.redd.it/pxp048081juh1.png?width=3200&format=png&auto=webp&s=5ee30548379fa9e7edce9e0242844656f2e89c07