My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context
A user is running local AI with two BC-250 ex-mining APUs, costing $115 each, connected via llama.cpp with Vulkan and RPC on Bazzite. These boards offer approximately 27GB of combined GPU memory and communicate over 1gb Ethernet. The setup, totaling around $300 including the PSU, achieves 60 tok/s with 64k context when running Qwen3.6-35B-A3B Q4_K_M. The user plans to expand to six boards to test Qwen 3.8 flash.
This report details a specific, low-cost hardware setup for local AI, unlike general discussions of AI models or software, and includes concrete performance metrics for Qwen3.6-35B-A3B.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 24, 2026, 12:02 UTC
- Ingested
- Sep 24, 2026, 12:02
- Source type
- Dev community
Full text isn't available here.
Read at source →