Skip to content
RCreddit.com·
Not on the current live radar

Is llama.cpp meant to be slow at long context, even when you aren't using that context?

AI summary

A user is experiencing slow performance with llama.cpp when using a Qwen 3.5 9B model with a 131K context, even when not fully utilizing the context. They are running it on an 8 GB laptop 4060 with 1GB reserved for the OS, using the command llama serve -hf bartowski/Ornith-1.5-9B-GGUF:IQ4_XS --fit on --cache-type-k q8_0 --cache-type-v q4_1 -c 131072 --temp 0.7. They are seeking an explanation for this slowdown compared to 16K context usage.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 23, 2026, 02:01 UTC

Ingested
Sep 23, 2026, 02:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com