RCreddit.com·
Not on the current live radar
This draft model is OP on 16 GB cards for Qwen 3.8 27b
A user reported that a draft model, specifically HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF, performs exceptionally well on 16 GB cards. When paired with ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF (IQ3_XXS with 128k context), it achieved an average token generation speed of about 60 tokens per second on an RX 9070 XT graphics card. Further testing by another individual on a different card is also mentioned.
This report uniquely details specific performance metrics for the HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF draft model on 16 GB cards, unlike general discussions of model efficiency.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 13, 2026, 06:01 UTC
- Ingested
- Sep 13, 2026, 06:01
- Source type
- Dev community
Full text isn't available here.
Read at source →