Skip to content
RCreddit.com·
Not on the current live radar

What's the best setup for Qwen3.8 27b for a 16 gig VRAM?

AI summary

A user on reddit.com inquired about the optimal setup for Qwen3.8 27b with 16 GB VRAM. They provided a llama.cpp command for llama-server, specifying parameters like --model ~/Documents/Models/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf, --host 0.0.0.0, --port 8001, --ngl 99, --flash-attn on, and --ctx-size 131072. The user reported achieving speeds of 35 tokens/second or more with this configuration.

Why this one

Unlike general discussions, this post provides a specific llama.cpp command and performance metrics (35+ t/s) for Qwen3.8 27b on 16GB VRAM, offering a concrete starting point.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 2, 2026, 03:00 UTC

Ingested
Oct 2, 2026, 03:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com