How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
一位开发者在双R9700配置上,使用940k tokens的KV缓存,实现了Qwen3.8 27B模型280 tok/s的性能。BetterBench测试结果显示,解码速度在JSON类别中达到280.0 tok/s,在散文类别中为116.4 tok/s。预填充测试中,2000 tokens提示词的预填充速度中位数为4695 PP t/s。该开发者已将其MXFP4镜像和代码库开源,展示了R9700s性能优化的显著进展。
2 Months ago I had made a post how I was working on my dual R9700's. It's wild to look back at where we were then and where things now stand.
Since then after many users commenting and complaining about developers doing the same thing. I threw out a discord link and expected maybe 5 other developers to join which I thought would be fun. The community has now grown to 1,200 users (mostly developers) and a ton of collaboration happening.
A few weeks ago I started working on building support for MXFP4 on top of DeadCode's radiance image. This made sense to me looking at the hardware and I was happy when I had hit parity on performance between MXFP4 and FP8. The MXFP4 kernels use W4A8 which was something new and we have now blown past the performance of FP8 and appears like this is now the hardware limits of these cards.
Qwen3.8 27B w/ DFlash2
BetterBench decode results for Qwen3.8 27B w/ DFlash2 category decode t/s step ms tok/update json 280.0 22.92 6.17 math 254.2 23.08 5.81 file_edit 250.1 23.03 5.54 code 226.3 23.01 5.17 reasoning 194.3 23.19 4.32 summarization 190.6 23.01 4.40 chat 148.3 22.82 3.33 prose 116.4 23.14 2.65 BetterBench Prefill Results target depth prompt tokens TTFT p50 PP t/s median 2000 1514 323 ms 4695 8000 5918 1.21 s 4894 16000 11794 2.47 s 4779 32000 23543 4.98 s 4729 64000 47056 10.8 s 4377 128000 94065 24.6 s 3831 250000 183678 59.1 s 3106
This has been so fun working on these R9700's and driving them to peak performance. My entire image and repo for MXFP4 is open source also: https://codeberg.org/ggz14/radiance-vllm-mxfp4