Skip to content
RCreddit.com·
Not on the current live radar

VLLM 4x rtx 3060 vs 8x rtx 3060 performance loss

AI summary

A user is seeking advice on expected performance when scaling GPU setups for local AI servers, specifically comparing VLLM with four NVIDIA RTX 3060 GPUs versus eight. Currently, they achieve 25 tps with Qwen 3.8 27b Q6 using three RTX 3060s in llama.cpp layered mode. They anticipate an upgrade to around 50 tps with four GPUs and are curious about the performance implications of moving to an eight-GPU configuration, aiming for optimal cost-effectiveness with their existing hardware.

Why this one

This user's query moves beyond typical performance questions by directly comparing VLLM scaling from four to eight RTX 3060 GPUs, unlike most discussions focusing on single-card or dual-card setups.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 25, 2026, 16:02 UTC

Ingested
Sep 25, 2026, 16:02
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com