VLLM 4x rtx 3060 vs 8x rtx 3060 performance loss
A user is seeking advice on expected performance when scaling GPU setups for local AI servers, specifically comparing VLLM with four NVIDIA RTX 3060 GPUs versus eight. Currently, they achieve 25 tps with Qwen 3.8 27b Q6 using three RTX 3060s in llama.cpp layered mode. They anticipate an upgrade to around 50 tps with four GPUs and are curious about the performance implications of moving to an eight-GPU configuration, aiming for optimal cost-effectiveness with their existing hardware.
This user's query moves beyond typical performance questions by directly comparing VLLM scaling from four to eight RTX 3060 GPUs, unlike most discussions focusing on single-card or dual-card setups.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月25日 16:02 UTC
- 收录
- 2026年9月25日 16:02
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →