RCreddit.com·
Not on the current live radar
... so, yeah.
A user successfully ran Qwen3.8-Flash-Next-GSQ-RCO-GGUF on an M4Pro 48GB Mac, initially noting 3.8-27B was faster. However, after applying specific llama.cpp flags like --flash-attn on and --lazy-mode on, Flash-Next outperformed 27B, handling up to 131K without OOM. The Q2_0 quantization of this model in llama.cpp demonstrated superior performance in prompt processing, speed generation, and intelligence compared to 27B at oQ4e.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 28, 2026, 00:00 UTC
- Ingested
- Sep 28, 2026, 00:00
- Source type
- Dev community
Full text isn't available here.
Read at source →