Skip to content
RCreddit.com·
Not on the current live radar

... so, yeah.

AI summary

A user successfully ran Qwen3.8-Flash-Next-GSQ-RCO-GGUF on an M4Pro 48GB Mac, initially noting 3.8-27B was faster. However, after applying specific llama.cpp flags like --flash-attn on and --lazy-mode on, Flash-Next outperformed 27B, handling up to 131K without OOM. The Q2_0 quantization of this model in llama.cpp demonstrated superior performance in prompt processing, speed generation, and intelligence compared to 27B at oQ4e.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 28, 2026, 00:00 UTC

Ingested
Sep 28, 2026, 00:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com