跳到正文
RCreddit.com·
暂不在当前实时榜单

Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

AI 摘要

A developer reported significant performance improvements for the Qwen3.8-Flash-Next model on a 12GB RTX 5070. Initially achieving 15 tok/s output and 100-120 tok/s prompt processing with IQ3_XXS quant using llama.cpp, they later developed a custom inference engine. This engine boosted performance to ~65 tok/s output and ~430 tok/s prompt processing with the same IQ3_XXS quant, with 2-bit quants using RCO-GSQ quantization running even faster.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月24日 20:01 UTC

收录
2026年9月24日 20:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com