返回
RCreddit.com
17
·20小时前·开发者社区 · RSS

Quad R9700 AI Pro with vLLM-Radiance easily reaching 17,6k PP

查看原文
Qwen

热度趋势

新上榜
最近 24 小时与此前 24 小时对比 · 7 天曲线

百分比基于当前可用热度信号,而非评论数或独立用户人数。

AI 摘要

一名用户报告称,通过vLLM-Radiance,其四路R9700 AI Pro配置轻松达到了17,6k PP。该系统配置包括一块技嘉MZ32-AR0 (Rev 1.0)主板和一颗EPYC 7282处理器。他们正在使用Hermes运行Qwen 3.8 27b fp8,上下文为262k。值得注意的是,其中一块GPU通过PCIe 4x8运行,而另外三块则通过完整的4x16运行。

https://preview.redd.it/74bmvel9b5nh1.png?width=1602&format=png&auto=webp&s=0d0c1adaa016a486ffd97c4c466e980dc611b139

I've only recently started looking deeper into vLLM after running llama.cpp for a good while.

Initially vLLM (official repo) was terribly slow on my four R9700s (tried that one with two as well), however after trying radiance everything changed.

That prefill spike was two agent profiles working on different tasks simultaneously (one is writing a yt-dlp dl/conversion workflow the other is auditing agents (profiles).

Best TG i've hit was 106 Tok/s with a 80% MTP 4 acceptance rate.

For reference, I'm running a Gigabyte MZ32-AR0 (Rev 1.0), EPYC 7282 and using Hermes with Qwen 3.8 27b fp8 262k ctx - worth noting that one GPU is actually only running by PCIe 4x8, three full 4x16.

On that note i'm also happy to say that vLLM-Radiance does work well with a quad setup in my case - nvtop consistently shows 100% usage of the four cards, officially only dual setups are supported.

I hope this doesn't count as a low effort post, i just had to share.

Quad R9700 AI Pro with vLLM-Radiance easily reaching 17,6k PP · BuzzRadr