Skip to content
RCreddit.com·
Not on the current live radar

Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

AI summary

A user tested Gufo 0.4.0 with Qwen3.8 27B UD-Q4_K_XL and a DFlash2 Q4_K_M draft model, achieving 70.56 tok/s for a single user and 123 tok/s with eight users on a Strix Halo device. The setup used Gufo's Podman image and benchmark script with specific settings like greedy decoding and 128 output tokens. The user plans to further evaluate output quality.

Why this one

This report provides a first-hand, detailed account of Gufo 0.4.0's performance with specific Qwen models, unlike other discussions that only cite headline figures.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 2, 2026, 03:00 UTC

Ingested
Oct 2, 2026, 03:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com