RCreddit.com·
Not on the current live radar
Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm
A user reported high performance with Qwen3.8-27B, achieving over 70 tok/s (or 160 tok/s concurrent) and 10k tok/s prefill speed using vanilla vllm. This was on 2x3090 GPUs or Ampere/higher with 48GB+ memory, supporting full context. This performance reportedly surpassed others on similar hardware with half the speed and context. Benchmarking involved various quantized models, vllm, sglang, and llama.cp, tested on systems with AMD EPYC 7532 or AMD Ryzen Threadripper PRO 3945WX CPUs.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 23, 2026, 02:01 UTC
- Ingested
- Sep 23, 2026, 02:01
- Source type
- Dev community
Full text isn't available here.
Read at source →