RCreddit.com·
暂不在当前实时榜单
Is llama.cpp meant to be slow at long context, even when you aren't using that context?
A user is experiencing slow performance with llama.cpp when using a Qwen 3.5 9B model with a 131K context, even when not fully utilizing the context. They are running it on an 8 GB laptop 4060 with 1GB reserved for the OS, using the command llama serve -hf bartowski/Ornith-1.5-9B-GGUF:IQ4_XS --fit on --cache-type-k q8_0 --cache-type-v q4_1 -c 131072 --temp 0.7. They are seeking an explanation for this slowdown compared to 16K context usage.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月23日 02:01 UTC
- 收录
- 2026年9月23日 02:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →