RCreddit.com·
Not on the current live radar
Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?
A user is experimenting with Qwen3.8-Flash-Next on a system with a 5090 GPU and 64GB RAM, using Llama.cpp. They are curious about the Flash-Next version's performance, especially since it doesn't require the entire model to fit into VRAM. The user is running the model with specific parameters, including --gpu-layers all, --ctx-size 64768, and --flash-attn on. They are surprised by the speed, which is better than expected, but suspect they might be missing something regarding RAM usage.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 25, 2026, 16:02 UTC
- Ingested
- Sep 25, 2026, 16:02
- Source type
- Dev community
Full text isn't available here.
Read at source →