2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s
A user is running Qwen3-Flash-Next (177B / ~6B active MoE, IQ4_XS) on Ilama.cpp with a system equipped with two RTX 3090 graphics cards and an EPYC processor. This configuration achieves approximately 38 tok/s in single-stream mode. However, when processing two parallel requests simultaneously, performance significantly drops to about 4 tok/s per request. The user aims to improve performance when multiple agents run in parallel.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月12日 16:52 UTC
收录当时偏移:UTC+02026年9月13日 15:01 UTC
- 发布
- 2026年9月12日 16:52
- 收录
- 2026年9月13日 15:01
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
What I have:
- CPU: EPYC 7551 (32c/64T, Zen 1)
- Board: Supermicro H11SSL-i (SP3), Rev 2.0
- RAM: 128 GB DDR4-2133 (all 8 channels full)
- GPU: 2x RTX 3090 (48 GB total, PCIe 3.0)
- 1500 W PSU
What I run:
- Qwen3-Flash-Next (177B total / ~6B active MoE, IQ4_XS) on Ilama.cpp. Experts live in system RAM, hot ones cached in VRAM. Single stream = 38 tok/s. Two parallel requests drop to ~4 tok/s each.
Budget:
~$800. Realistically that's either one more RTX 3090 or a CPU upgrade (a Zen 2 "Rome" EPYC drops into the same board). A new motherboard is out of budget i think for now.
Which gives more inference speed for this setup - adding the 3rd 3090, or swapping to a faster/newer CPU?
And would more/faster RAM matter here? Curious what people running similar rigs have actually measured.
I am also interested in having multiple agents running at the same time, which currently slows it down heavily, so keeping the performance at multiple agents parallel would be a huge boost as well!