Back
RCreddit.com
13
·1 days ago·Dev community · RSS

Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS

View original
QwenModel release

Heat trend

New
Latest 24h versus previous 24h · 7-day curve

The percentage is based on available heat signal, not comment count or independent people.

Why it matters

Qwen model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.

AI summary

A user on Reddit reported achieving over 200k context on 16GB VRAM using Qwen 3.8 27B UD-IQ3_XXS. Previously, they used UD-Q3_K_XL with over 140000 context, noting its good quality and few erroneous tool calls. While the new IQ3_XXS model allows for higher context, prompt processing speed decreased from 700-800 tk/s to 400 tk/s, and quality differences are still being evaluated.

I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others here reporting good results with IQ3_XXS, so I gave it a try.

The downside is prompt processing speed went down from 700-800 tk/s to 400 tk/s. Quality difference is yet to be tested.

Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS · BuzzRadr