Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3)
A user successfully deployed Qwen 3.8 Next on a v100 6GPU setup, configured with TP2 PP3. Performance metrics show varying token generation speeds across different input lengths, ranging from 1,389 tok/s for 1K tokens to 2,759 tok/s for 131K tokens. The system achieved a maximum speed of 4,679 tok/s at 16,384 (16K) input length. The user expressed satisfaction with the system's stability, thermal management, and low noise levels.
This report is the first to detail Qwen 3.8 Next's performance on a v100 6GPU setup, providing specific token generation speeds across various input lengths, unlike general benchmarks.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 20, 2026, 11:01 UTC
- Ingested
- Sep 20, 2026, 11:01
- Source type
- Dev community
Full text isn't available here.
Read at source →