Skip to content
RCreddit.com·
Not on the current live radar

Running Qwen3.8-27B-Q4 at max context on a 32 GB GPU while avoiding kvcache quantization

AI summary

A user successfully ran the Qwen3.8-27B-Q4 model at a maximum context of 262,144 tokens on a 32 GB GPU, while avoiding kvcache quantization. The baseline achieved 55,040 tokens with 17.13 GiB for weights and 5.65 GiB for context. Strategies like moving mmproj-to-cpu and disabling-spec significantly increased context, with the final strategy involving quantize-kv-q4 to reach the maximum context.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 12, 2026, 06:01 UTC

Ingested
Sep 12, 2026, 06:01
Source type
Dev community
Article

Full text isn't available here.

Read at source →
Source·reddit.com