Quad R9700 AI Pro with vLLM-Radiance easily reaching 17,6k PP
Heat trend
The percentage is based on available heat signal, not comment count or independent people.
A user reported that their quad R9700 AI Pro setup, utilizing vLLM-Radiance, effortlessly achieved 17,6k PP. The system configuration includes a Gigabyte MZ32-AR0 (Rev 1.0) motherboard and an EPYC 7282 processor. They are running Qwen 3.8 27b fp8 with Hermes, using a 262k context. Notably, one GPU operates via PCIe 4x8, while the other three run on full 4x16.
https://preview.redd.it/74bmvel9b5nh1.png?width=1602&format=png&auto=webp&s=0d0c1adaa016a486ffd97c4c466e980dc611b139
I've only recently started looking deeper into vLLM after running llama.cpp for a good while.
Initially vLLM (official repo) was terribly slow on my four R9700s (tried that one with two as well), however after trying radiance everything changed.
That prefill spike was two agent profiles working on different tasks simultaneously (one is writing a yt-dlp dl/conversion workflow the other is auditing agents (profiles).
Best TG i've hit was 106 Tok/s with a 80% MTP 4 acceptance rate.
For reference, I'm running a Gigabyte MZ32-AR0 (Rev 1.0), EPYC 7282 and using Hermes with Qwen 3.8 27b fp8 262k ctx - worth noting that one GPU is actually only running by PCIe 4x8, three full 4x16.
On that note i'm also happy to say that vLLM-Radiance does work well with a quad setup in my case - nvtop consistently shows 100% usage of the four cards, officially only dual setups are supported.
I hope this doesn't count as a low effort post, i just had to share.