Skip to content
RCreddit.com·

Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards

AI summary

A new release of Qwen3.8 27b is now available, specifically optimized for 12GB Ampere cards. This version introduces configurable runtime refinements, including compact MTP caches and 16-bit activations, alongside a new Staged + Journaled KVaRN variant that reduces KLD by approximately 40%. Users can achieve context lengths of 205k to 230k with 11GB of memory, with further increases possible in headless configurations.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Oct 9, 2026, 21:39 UTC

IngestedOffset at this time: UTC+0Oct 10, 2026, 00:00 UTC

Published
Oct 9, 2026, 21:39
Ingested
Oct 10, 2026, 00:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

No comparison yet
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

But, it is mostly targeted to the 12GB card crowd, who had not been the focus of the previous two releases. In addition to a series of configurable runtime refinements + adjustments (compact MTP caches, 16 bit activations, redundant overhead items removed), I've also added a new variant of KVaRN (Staged + Journaled KVaRN), that reduced KLD vs paper faithful + competitive versions by ~40%. 4/4 is my new recommended default (significantly better performance vs a 8/4 KV cache) for larger cards, but for people looking to get maximum ctx, the 3/3 bit (0.001 nats KLD) and 3/2 (0.0024 nats) are still very solid (the KLD from 3/2 is smaller than the performance drop you see going from 4 XL to 4 S quants). They support ctx lengths of 205k to 230k for the model tested, respectively, at 11 GB (significantly more if you're using 100% of GPU in a headless config).

The model tested was 2.3bpw fusion of swift-1.5-uncensored and mirai's 2.5 bpw model. It retained 85% of the BF16's performance on LiveCode Bench across 7 runs on 3060/80/ti cards. On 3080/ti it averaged ~65-70 tps (thanks to MTP) on the evals.

By no means is the model lossless, and I won't pretend that it is like a certain other model did. But using it personally in hermes for a couple days, and running it through coding evals + pi testing, it is genuinely usable for standard agentic tasks. Subjectively, I'd put it somewhere between the BF16 versions of qwen3.5 and 3.6, which is pretty good for a model that fits in 12GB with context!

Models (4 XS-M, 2.3bpw, etc) based on swift 1.5 are here: https://huggingface.co/collections/jakeatx/qwen38-27b-models (llamAmpere supports EXL if you prefer that family, but I am working on improving prefill + decode kernels, so still not quite first class speeds yet in the 3-5 bpw range)

As before, please share your results, bugs, etc. They will get added to the backlog.

Source·reddit.com