跳到正文
RCreddit.com·

Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM

AI 摘要

A user successfully ran Qwen 3.6 35B A3B with a 131K context window and vision capabilities on an RTX 2060 6GB GPU. Utilizing llama.cpp, they achieved approximately 600 tokens/second prefill and 23 tokens/second decode speeds. The setup involved the HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M model and a Q8 KV cache, demonstrating impressive performance on limited VRAM.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月10日 05:41 UTC

收录当时偏移:UTC+02026年10月10日 14:00 UTC

发布
2026年10月10日 05:41
收录
2026年10月10日 14:00
来源类型
开发者社区
档位
社区
信源状态
同步延迟

档位是按信源手工设定的编辑判断,不是逐条打分。

Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M with llama.cpp at ~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.

Speeds start at ~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.

The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under ~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys ~1GB of VRAM.

Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.

Launch command:

bat

@echo off

cd /d "%~dp0"

"%~dp0llama-server.exe" ^

-m "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf" ^

--mmproj "C:\Qwen3.6-35B-A3B-Uncensored-Q4_K_M\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" ^

--no-mmproj-offload ^

-ngl 99 ^

--n-cpu-moe 39 ^

-c 131072 ^

-np 1 ^

-t 6 ^

-tb 10 ^

-b 2048 ^

-ub 2048 ^

-fa on ^

-ctk q8_0 ^

-ctv q8_0 ^

--load-mode mmap+mlock ^

来源·reddit.com