Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
A user successfully ran the Qwen3.8 Flash Next 176B MoE model on a laptop equipped with a 16GB RTX 3080, 32GB RAM, and an SSD. The experiment aimed to test the limits of a standard laptop with a large model. Performance metrics showed a decode throughput of 11.09 tokens/s and a whole-process time of 16.54s, with a GPU peak of 14,832.5 MiB and an OS peak working set of 18.51 GiB. The user is seeking comparisons with other implementations like llama.cpp or Strata.
This report details the first successful attempt to run the Qwen3.8 Flash Next 176B MoE model on a consumer-grade laptop, unlike previous tests on more powerful hardware.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 4, 2026, 00:00 UTC
- Ingested
- Oct 4, 2026, 00:00
- Source type
- Dev community
Full text isn't available here.
Read at source →