跳到正文
RCreddit.com·
暂不在当前实时榜单

Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.

AI 摘要

A user compared the performance of Qwen3.8-Flash-Next across llama.cpp, SGLang, and FreeToken, noting a significant difference in time to first token at full context (35s vs 258s). The findings also detailed the impact of MTP (Multi-Threaded Processing) on VRAM usage and token generation rates. At 96 GiB, MTP provided 1.61x faster decode, but at 24 GiB, it made decode about 3.4x slower. The user is interested in further testing on different GPU/memory setups and MTP's effectiveness with CPU offloading.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月8日 22:00 UTC

收录
2026年9月8日 22:00
来源类型
开发者社区
正文

本站未收录正文。

前往源站阅读 →
来源·reddit.com