VLLM 4x rtx 3060 vs 8x rtx 3060 performance loss
A user is seeking advice on expected performance when scaling GPU setups for local AI servers, specifically comparing VLLM with four NVIDIA RTX 3060 GPUs versus eight. Currently, they achieve 25 tps with Qwen 3.8 27b Q6 using three RTX 3060s in llama.cpp layered mode. They anticipate an upgrade to around 50 tps with four GPUs and are curious about the performance implications of moving to an eight-GPU configuration, aiming for optimal cost-effectiveness with their existing hardware.
This user's query moves beyond typical performance questions by directly comparing VLLM scaling from four to eight RTX 3060 GPUs, unlike most discussions focusing on single-card or dual-card setups.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 25, 2026, 16:02 UTC
- Ingested
- Sep 25, 2026, 16:02
- Source type
- Dev community
Full text isn't available here.
Read at source →