Skip to content
RCreddit.com·
Not on the current live radar

2x CMP 170HX 64GB: GLM-5.3-Flash at 384K context / ~90 tok/s (EXL3, HBM-first setup) + Qwen3.8 comparison

AI summary

A user has successfully configured two 64GB CMP 170HX cards to run GLM-5.3-Flash, achieving a stable setup with 384K context and approximately 90 tokens per second using EXL3. The user's target quantization size is EXL3 3.05bpw, estimated to provide around 93.05% top-1 agreement. Future plans include attempting to reach 1M context, acknowledging that prefill time will likely be a significant challenge.

Why this one

This report details a specific GLM-5.3-Flash configuration on dual CMP 170HX cards, unlike general discussions, and includes a Qwen3.8 comparison.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 8, 2026, 03:00 UTC

Ingested
Oct 8, 2026, 03:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com