Back
RCreddit.com
12
·2 days ago·RSS
Not on the current live radar

Experience report - Qwen 3.8 Flash Next on memory rich, GPU poor setup

View original
LlamaQwenModel release

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

Why it matters

Llama model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.

AI summary

A user shared an experience report on running Qwen 3.8 Flash Next on a memory-rich, GPU-poor setup. The setup utilized llama-server with a Qwen3.6-35B-A3B model, specifying --ctx-size 131744, --batch-size 1024, and --flash-attn on. The user also configured --n-gpu-layers 999 and --n-cpu-moe 26, seeking comparisons with other constrained setups.