Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm

AI summary

A user reported high performance with Qwen3.8-27B, achieving over 70 tok/s (or 160 tok/s concurrent) and 10k tok/s prefill speed using vanilla vllm. This was on 2x3090 GPUs or Ampere/higher with 48GB+ memory, supporting full context. This performance reportedly surpassed others on similar hardware with half the speed and context. Benchmarking involved various quantized models, vllm, sglang, and llama.cp, tested on systems with AMD EPYC 7532 or AMD Ryzen Threadripper PRO 3945WX CPUs.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 23, 2026, 02:01 UTC

Ingested
Sep 23, 2026, 02:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com