Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?

AI summary

A user is experimenting with Qwen3.8-Flash-Next on a system with a 5090 GPU and 64GB RAM, using Llama.cpp. They are curious about the Flash-Next version's performance, especially since it doesn't require the entire model to fit into VRAM. The user is running the model with specific parameters, including --gpu-layers all, --ctx-size 64768, and --flash-attn on. They are surprised by the speed, which is better than expected, but suspect they might be missing something regarding RAM usage.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 25, 2026, 16:02 UTC

Ingested
Sep 25, 2026, 16:02
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com