600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"?
A user achieved 600 tokens/second on a single request using the qwen3.6 35ba3b model with Ninfer on an RTX Pro 6000. While acknowledging it's not the most intelligent model, they find it suitable for "read+find" or coding tasks where brute-force processing is acceptable. Despite potentially using 20 times more tokens, its speed surpasses many other local models, making it a "stupid fast" and enjoyable tool.
This report highlights a specific model, qwen3.6 35ba3b, achieving 600 tokens/second, a speed described as "stupid fast" compared to many other local models.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 18, 2026, 07:00 UTC
- Ingested
- Sep 18, 2026, 07:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →