Skip to content
RCreddit.com·
Not on the current live radar

NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

AI summary

A developer has achieved 400 tokens/second (tg/s) with the Qwen 3.8 Flash Next model on an RTX6000 using a modified NInfer fork. Benchmarks show prompt processing speeds up to 13,908 tokens/second (pp/s) for an 8K prompt length. Performance varies with prompt length and bit precision, with 16-bit generally outperforming 8-bit for longer prompts. Further details and benchmarks are available in the project's GitHub repository.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 7, 2026, 06:00 UTC

Ingested
Oct 7, 2026, 06:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com