跳到正文
RCreddit.com·
暂不在当前实时榜单

Is llama.cpp meant to be slow at long context, even when you aren't using that context?

AI 摘要

A user is experiencing slow performance with llama.cpp when using a Qwen 3.5 9B model with a 131K context, even when not fully utilizing the context. They are running it on an 8 GB laptop 4060 with 1GB reserved for the OS, using the command llama serve -hf bartowski/Ornith-1.5-9B-GGUF:IQ4_XS --fit on --cache-type-k q8_0 --cache-type-v q4_1 -c 131072 --temp 0.7. They are seeking an explanation for this slowdown compared to 16K context usage.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月23日 02:01 UTC

收录
2026年9月23日 02:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com