跳到正文
RCreddit.com·

Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide

AI 摘要

A user successfully ran Qwen 3.8 27B Q4 XS with 100K context on a 16GB AMD RX 7800 XT GPU, achieving approximately 30 t/s decode speed. This setup, which many believed infeasible on 16GB VRAM, utilized specific llama-server parameters. Key configurations included --n-gpu-layers 999, --ctx-size 100096, --cache-type-k q8_0, and --cache-type-v q5_1 to optimize performance and memory usage for the Qwen3.8-27B-UD-IQ4_XS.gguf model.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年9月29日 11:18 UTC

收录当时偏移:UTC+02026年9月29日 20:00 UTC

发布
2026年9月29日 11:18
收录
2026年9月29日 20:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

I'm running Qwen 3.8 27B Q4 XS with ~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.

Build llama.cpp with Vulkan:

cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j

Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:

llama-server \ --model Qwen3.8-27B-UD-IQ4_XS.gguf \ --mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \ --n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \ --batch-size 2048 --ubatch-size 512 \ --flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \ --load-mode none --fit off \ --cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \ --no-context-shift --jinja --reasoning-format deepseek \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \ --threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080

来源·reddit.com