RCreddit.com·
Not on the current live radar
2x CMP 170HX 64GB: GLM-5.3-Flash at 384K context / ~90 tok/s (EXL3, HBM-first setup) + Qwen3.8 comparison
A user has successfully configured two 64GB CMP 170HX cards to run GLM-5.3-Flash, achieving a stable setup with 384K context and approximately 90 tokens per second using EXL3. The user's target quantization size is EXL3 3.05bpw, estimated to provide around 93.05% top-1 agreement. Future plans include attempting to reach 1M context, acknowledging that prefill time will likely be a significant challenge.
This report details a specific GLM-5.3-Flash configuration on dual CMP 170HX cards, unlike general discussions, and includes a Qwen3.8 comparison.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 8, 2026, 03:00 UTC
- Ingested
- Oct 8, 2026, 03:00
- Source type
- Dev community
Full text isn't available here.
Read at source →