RCreddit.com·
暂不在当前实时榜单
Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm
A user reported high performance with Qwen3.8-27B, achieving over 70 tok/s (or 160 tok/s concurrent) and 10k tok/s prefill speed using vanilla vllm. This was on 2x3090 GPUs or Ampere/higher with 48GB+ memory, supporting full context. This performance reportedly surpassed others on similar hardware with half the speed and context. Benchmarking involved various quantized models, vllm, sglang, and llama.cp, tested on systems with AMD EPYC 7532 or AMD Ryzen Threadripper PRO 3945WX CPUs.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月23日 02:01 UTC
- 收录
- 2026年9月23日 02:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →