Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards
A new release of Qwen3.8 27b is now available, specifically optimized for 12GB Ampere cards. This version introduces configurable runtime refinements, including compact MTP caches and 16-bit activations, alongside a new Staged + Journaled KVaRN variant that reduces KLD by approximately 40%. Users can achieve context lengths of 205k to 230k with 11GB of memory, with further increases possible in headless configurations.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月9日 21:39 UTC
收录当时偏移:UTC+02026年10月10日 00:00 UTC
- 发布
- 2026年10月9日 21:39
- 收录
- 2026年10月10日 00:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 同步延迟
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
But, it is mostly targeted to the 12GB card crowd, who had not been the focus of the previous two releases. In addition to a series of configurable runtime refinements + adjustments (compact MTP caches, 16 bit activations, redundant overhead items removed), I've also added a new variant of KVaRN (Staged + Journaled KVaRN), that reduced KLD vs paper faithful + competitive versions by ~40%. 4/4 is my new recommended default (significantly better performance vs a 8/4 KV cache) for larger cards, but for people looking to get maximum ctx, the 3/3 bit (0.001 nats KLD) and 3/2 (0.0024 nats) are still very solid (the KLD from 3/2 is smaller than the performance drop you see going from 4 XL to 4 S quants). They support ctx lengths of 205k to 230k for the model tested, respectively, at 11 GB (significantly more if you're using 100% of GPU in a headless config).
The model tested was 2.3bpw fusion of swift-1.5-uncensored and mirai's 2.5 bpw model. It retained 85% of the BF16's performance on LiveCode Bench across 7 runs on 3060/80/ti cards. On 3080/ti it averaged ~65-70 tps (thanks to MTP) on the evals.
By no means is the model lossless, and I won't pretend that it is like a certain other model did. But using it personally in hermes for a couple days, and running it through coding evals + pi testing, it is genuinely usable for standard agentic tasks. Subjectively, I'd put it somewhere between the BF16 versions of qwen3.5 and 3.6, which is pretty good for a model that fits in 12GB with context!
Models (4 XS-M, 2.3bpw, etc) based on swift 1.5 are here: https://huggingface.co/collections/jakeatx/qwen38-27b-models (llamAmpere supports EXL if you prefer that family, but I am working on improving prefill + decode kernels, so still not quite first class speeds yet in the 3-5 bpw range)
As before, please share your results, bugs, etc. They will get added to the backlog.