Back
RCreddit.com
18
·17 hr ago·Dev community · RSS

How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache

View original

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

AI summary

A developer achieved 280 tok/s on Qwen3.8 27B using dual R9700s and a 940k tokens KV cache. This performance was measured with BetterBench, showing decode speeds up to 280.0 tok/s for JSON and 116.4 tok/s for prose. Prefill results indicate a median prefill speed of 4695 PP t/s for a 2000-token prompt. The developer open-sourced their MXFP4 image and repository, highlighting significant progress since their initial work on the R9700s.

2 Months ago I had made a post how I was working on my dual R9700's. It's wild to look back at where we were then and where things now stand.

Since then after many users commenting and complaining about developers doing the same thing. I threw out a discord link and expected maybe 5 other developers to join which I thought would be fun. The community has now grown to 1,200 users (mostly developers) and a ton of collaboration happening.

A few weeks ago I started working on building support for MXFP4 on top of DeadCode's radiance image. This made sense to me looking at the hardware and I was happy when I had hit parity on performance between MXFP4 and FP8. The MXFP4 kernels use W4A8 which was something new and we have now blown past the performance of FP8 and appears like this is now the hardware limits of these cards.

Qwen3.8 27B w/ DFlash2

BetterBench decode results for Qwen3.8 27B w/ DFlash2 category decode t/s step ms tok/update json 280.0 22.92 6.17 math 254.2 23.08 5.81 file_edit 250.1 23.03 5.54 code 226.3 23.01 5.17 reasoning 194.3 23.19 4.32 summarization 190.6 23.01 4.40 chat 148.3 22.82 3.33 prose 116.4 23.14 2.65 BetterBench Prefill Results target depth prompt tokens TTFT p50 PP t/s median 2000 1514 323 ms 4695 8000 5918 1.21 s 4894 16000 11794 2.47 s 4779 32000 23543 4.98 s 4729 64000 47056 10.8 s 4377 128000 94065 24.6 s 3831 250000 183678 59.1 s 3106

This has been so fun working on these R9700's and driving them to peak performance. My entire image and repo for MXFP4 is open source also: https://codeberg.org/ggz14/radiance-vllm-mxfp4

How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache · BuzzRadr