Skip to content
RCreddit.com·

Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide

AI summary

A user successfully ran Qwen 3.8 27B Q4 XS with 100K context on a 16GB AMD RX 7800 XT GPU, achieving approximately 30 t/s decode speed. This setup, which many believed infeasible on 16GB VRAM, utilized specific llama-server parameters. Key configurations included --n-gpu-layers 999, --ctx-size 100096, --cache-type-k q8_0, and --cache-type-v q5_1 to optimize performance and memory usage for the Qwen3.8-27B-UD-IQ4_XS.gguf model.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 29, 2026, 11:18 UTC

IngestedOffset at this time: UTC+0Sep 29, 2026, 20:00 UTC

Published
Sep 29, 2026, 11:18
Ingested
Sep 29, 2026, 20:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

I'm running Qwen 3.8 27B Q4 XS with ~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.

Build llama.cpp with Vulkan:

cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j

Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:

llama-server \ --model Qwen3.8-27B-UD-IQ4_XS.gguf \ --mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \ --n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \ --batch-size 2048 --ubatch-size 512 \ --flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \ --load-mode none --fit off \ --cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \ --no-context-shift --jinja --reasoning-format deepseek \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \ --threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080

Source·reddit.com