Skip to content
RCreddit.com·
Not on the current live radar

600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"?

AI summary

A user achieved 600 tokens/second on a single request using the qwen3.6 35ba3b model with Ninfer on an RTX Pro 6000. While acknowledging it's not the most intelligent model, they find it suitable for "read+find" or coding tasks where brute-force processing is acceptable. Despite potentially using 20 times more tokens, its speed surpasses many other local models, making it a "stupid fast" and enjoyable tool.

Why this one

This report highlights a specific model, qwen3.6 35ba3b, achieving 600 tokens/second, a speed described as "stupid fast" compared to many other local models.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 18, 2026, 07:00 UTC

Ingested
Sep 18, 2026, 07:00
Source type
Dev community

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Full text isn't available here.

Read at source →
Source·reddit.com