Skip to content
RCreddit.com·
Not on the current live radar

7900 XTX — two "low-thinking" Qwen 3.8 27B quants (Swift + ThinkingCap) vs the regular quant

AI summary

Performance tests on a 7900 XTX compared two "low-thinking" Qwen 3.8 27B quants, Swift and ThinkingCap, against a regular quant. The regular quant (unsloth) achieved ~530 t/s prefill and ~48 t/s decode, with a runtime of ~1407s. ThinkingCap showed ~100 t/s prefill and ~43 t/s decode over ~1087s, while Swift managed ~614 t/s prefill and ~32 t/s decode over ~1400s. The prefill speed of ThinkingCap was noted as unusually low, suggesting further testing is needed.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 25, 2026, 16:02 UTC

Ingested
Sep 25, 2026, 16:02
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com