Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-27B with just 12GB VRAM + 8GB RAM - ~18 tok/sec

AI summary

A developer successfully ran Qwen3.8-27B at approximately 18 tokens/second decode and 500-600 tokens/second prefill with a 64k context, using only 12GB VRAM and 8GB RAM. This was achieved using the ISTA-DASLab Qwen3.8-27B-GSQ-RCO IQ3_S quant, which consumed about 11GB. The same recipe also worked with 0bserverx’ Qwen3.8-27B-Heretic-GSQ-RCO IQ3_S quant, showing similar performance with a slight speed reduction. The full build and serve scripts are available on GitHub for reproduction.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 7, 2026, 20:00 UTC

Ingested
Oct 7, 2026, 20:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com