跳到正文
HNHacker News·
暂不在当前实时榜单

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

AI 摘要

The Qwen 3.8 Flash Next (125B) model can run on consumer hardware like the RTX 4090 at speeds up to 100 T/s. Performance metrics for different quantization levels (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, Coder) on NVIDIA and AMD GPUs detail tokens per second for answer generation and prompt reading. For instance, an RTX 3090 (24 GB) is expected to achieve 100-140 tokens per second. This is enabled by the open-source Strata engine. A detailed performance table is available in DETAILS.md, and GPUs with more VRAM generally offer faster performance.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月4日 14:00 UTC

收录
2026年10月4日 14:00
来源类型
开发者社区
爆款判定
判定依据
热度约为该来源近期上榜条目中位水平的 6.6 倍
指标对比
227 vs 中位 34.5(20 条基线样本)
检出时间
10/04 14:00

本站未收录正文。

前往源站阅读 →
来源·Hacker News·github.com