Back
RCreddit.com

Best way to run Qwen3.8-27B on a system with a RTX 5090 + RTX 5070 Ti (32GB + 16GB)?

NVIDIAOn-device
Time & source
Published
09/05, 17:34
Ingested
09/06, 16:00
Source type
Dev community
Tier
Community
Source status
Healthy
Tier is a per-source editorial setting, not a per-item score.

I have a system with 2 GPUs and 48GB VRAM total, a RTX 5090 + RTX 5070Ti.

What would you say is the best way to run Qwen3.8-27B on that system with the best quality and 262k context?

Would just the normal llama.cpp work with how it detects and does its own magic with dual CPU systems, or something else?

I think the RTX5090 has pcie4 x16 and the RTX5070Ti has pcie x8 if that matters.