Skip to content
RCreddit.com·

Qwen3.8-27B UD-IQ4_XS Heretic + MTP on a 16 GB card with 55-68 tok/s (24gb and 12gb versions available too)

AI summary

A user successfully ran the uncensored Qwen3.8-27B (llmfan46's Heretic build, MTP head preserved) on an RTX 4080. The UD-IQ4_XS quantization, using 16 GB, achieved 55 tok/s on code and 50 tok/s on prose with MTP enabled, significantly faster than the 29 tok/s without MTP. Other versions, including 24 GB and 12 GB, are also available, with varying speeds and quality trade-offs.

Why this one

This report details a specific quantization of Qwen3.8-27B that achieves 55-68 tok/s on a 16 GB card, unlike other versions that offer different speed-quality trade-offs.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Oct 9, 2026, 22:54 UTC

IngestedOffset at this time: UTC+0Oct 10, 2026, 07:00 UTC

Published
Oct 9, 2026, 22:54
Ingested
Oct 10, 2026, 07:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

No comparison yet
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

I wanted the uncensored Qwen3.8-27B (llmfan46's Heretic build, MTP head preserved) on my RTX 4080 with MTP on (because I've been testing TONS of configs, and realized that would be the one)

And none of the published quants were built for that: the good IQ4_XS doesn't leave room for MTP, and the ones that fit are 3-bit.

I was amazed by that. as 16gb is VERY common, so decided to quantize it to fit, work well and loose as less as possible in quality:

I copied Unsloth's per-tensor UD recipe onto llmfan46's BF16, made 12 / 16 / 24 GB versions, and measured KLD against a Q8_0 of the same weights for every quant I could find.

Quality (mean KLD vs Q8_0, wikitext-2 / llama.cpp source code, lower is better)

Quant GiB Prose Code UD-Q5_K_XL (mine, 24 GB) 19.44 0.0045 0.0037 mradermacher i1-IQ4_XS 14.26 0.0202 0.0149 UD-IQ4_XS (mine, 16 GB) 13.27 0.0268 0.0192 llmfan46 Q3_K_M 13.48 0.0639 0.0456 mradermacher i1-IQ3_M 11.89 0.0649 0.0465 UD-IQ3_XXS (mine, 12 GB) 10.18 0.0904 0.0587 mradermacher i1-Q2_K 10.12 0.1551 0.1017 Speed on the 4080 (llama.cpp b11457, 15K-token prompt, 32K ctx, MTP 2 drafts, mean of 3 fixed seeds): UD-IQ4_XS does 55 tok/s on code and 50 on prose, against 29 / 29 without MTP. i1-IQ3_M is faster (70 / 54) at 2.4× the KLD, I prefer quality here over speed, but your call.

- 2 draft tokens beat 3. On the same seeds at 40K: 68 vs 56 tok/s on code, 55 vs 36 on prose. Fewer rejected drafts.

- 48K vs 40K produced the exact same tokens (identical draft-acceptance counts per seed) and was 22% slower. That's pure VRAM spill into shared memory, nothing else.

- The 16 GB edge is brutal. The same 40K config gave 68, 51 and 42 tok/s on code depending only on whether the desktop was holding 1.2, 1.4 or 1.9 GB of VRAM. Opening a Chrome window mid-generation took decode to 0.2 tok/s. If you're on Windows with a monitor on the same card, watch "shared GPU memory" in Task Manager.

- MTP head precision barely matters. q6_K head vs IQ3_S head: 83% vs 83% acceptance on code, 58% vs 54% on prose.

Repo with all three files, the commands, and the scripts (recipe extraction, dry-run verification, KLD, seeded speed bench): https://huggingface.co/codavidgarcia/Qwen3.8-27B-Uncensored-Heretic-MTP-UD-GGUF

The 12 GB and 24 GB context numbers are computed from llama.cpp's reported buffers, not measured on those cards. If you run them, let me know what you get!

Credit to llmfan46 for the weights, Unsloth for the recipes and mradermacher for the imatrix

edit: added the KLD chart since a few people asked about the numbers, (lower is better)

https://preview.redd.it/pxp048081juh1.png?width=3200&format=png&auto=webp&s=5ee30548379fa9e7edce9e0242844656f2e89c07

Source·reddit.com