Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM
A user successfully ran Qwen 3.6 35B A3B with a 131K context window and vision capabilities on an RTX 2060 6GB GPU. Utilizing llama.cpp, they achieved approximately 600 tokens/second prefill and 23 tokens/second decode speeds. The setup involved the HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M model and a Q8 KV cache, demonstrating impressive performance on limited VRAM.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Oct 10, 2026, 05:41 UTC
IngestedOffset at this time: UTC+0Oct 10, 2026, 14:00 UTC
- Published
- Oct 10, 2026, 05:41
- Ingested
- Oct 10, 2026, 14:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M with llama.cpp at ~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.
Speeds start at ~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.
The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under ~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys ~1GB of VRAM.
Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.
Launch command:
bat
@echo off
cd /d "%~dp0"
"%~dp0llama-server.exe" ^
-m "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf" ^
--mmproj "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" ^
--no-mmproj-offload ^
-ngl 99 ^
--n-cpu-moe 39 ^
-c 131072 ^
-np 1 ^
-t 6 ^
-tb 10 ^
-b 2048 ^
-ub 2048 ^
-fa on ^
-ctk q8_0 ^
-ctv q8_0 ^
--load-mode mmap+mlock ^