RCreddit.com·
暂不在当前实时榜单
2x CMP 170HX 64GB: GLM-5.3-Flash at 384K context / ~90 tok/s (EXL3, HBM-first setup) + Qwen3.8 comparison
A user has successfully configured two 64GB CMP 170HX cards to run GLM-5.3-Flash, achieving a stable setup with 384K context and approximately 90 tokens per second using EXL3. The user's target quantization size is EXL3 3.05bpw, estimated to provide around 93.05% top-1 agreement. Future plans include attempting to reach 1M context, acknowledging that prefill time will likely be a significant challenge.
This report details a specific GLM-5.3-Flash configuration on dual CMP 170HX cards, unlike general discussions, and includes a Qwen3.8 comparison.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月8日 03:00 UTC
- 收录
- 2026年10月8日 03:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →