Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS
Heat trend
The percentage is based on available heat signal, not comment count or independent people.
Qwen model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.
A user on Reddit reported achieving over 200k context on 16GB VRAM using Qwen 3.8 27B UD-IQ3_XXS. Previously, they used UD-Q3_K_XL with over 140000 context, noting its good quality and few erroneous tool calls. While the new IQ3_XXS model allows for higher context, prompt processing speed decreased from 700-800 tk/s to 400 tk/s, and quality differences are still being evaluated.
I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others here reporting good results with IQ3_XXS, so I gave it a try.
The downside is prompt processing speed went down from 700-800 tk/s to 400 tk/s. Quality difference is yet to be tested.