Running Qwen 3.8 next on 16vram+32ram - A useful/fun post for the gpu poors
A Reddit user shared their experience running Qwen 3.8 Next on a system with 16GB VRAM and 32GB RAM, a challenging setup for the large model. They noted that the model's components, including the N-gram/PLE Embedding (~29.48 GB) and MoE Routed Experts (~34.89 GB), require significant memory. Running the model with mmap enabled in llama.cpp resulted in slow speeds of ~2 tok/sec, making it effectively useless. The user suggests that systems with 64GB RAM would benefit more from an optimized llama.cpp fork.
This report details a specific, challenging setup for Qwen 3.8 Next on limited hardware, unlike other guides that assume more robust systems.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 14, 2026, 12:00 UTC
- Ingested
- Sep 14, 2026, 12:00
- Source type
- Dev community
Full text isn't available here.
Read at source →