跳到正文
RCreddit.com·
暂不在当前实时榜单

NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

AI 摘要

A developer has achieved 400 tokens/second (tg/s) with the Qwen 3.8 Flash Next model on an RTX6000 using a modified NInfer fork. Benchmarks show prompt processing speeds up to 13,908 tokens/second (pp/s) for an 8K prompt length. Performance varies with prompt length and bit precision, with 16-bit generally outperforming 8-bit for longer prompts. Further details and benchmarks are available in the project's GitHub repository.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月7日 06:00 UTC

收录
2026年10月7日 06:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com