跳到正文
RCreddit.com·
暂不在当前实时榜单

Running Qwen3.8-27B-Q4 at max context on a 32 GB GPU while avoiding kvcache quantization

AI 摘要

A user successfully ran the Qwen3.8-27B-Q4 model at a maximum context of 262,144 tokens on a 32 GB GPU, while avoiding kvcache quantization. The baseline achieved 55,040 tokens with 17.13 GiB for weights and 5.65 GiB for context. Strategies like moving mmproj-to-cpu and disabling-spec significantly increased context, with the final strategy involving quantize-kv-q4 to reach the maximum context.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月12日 06:01 UTC

收录
2026年9月12日 06:01
来源类型
开发者社区
正文

本站未收录正文。

前往源站阅读 →
来源·reddit.com