I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window.
A developer successfully used an iPhone as a secondary GPU for a 24 GB MacBook, enhancing the performance of Qwen 3.8 27B. This setup resulted in 29–44% faster prefilling and allowed the iPhone to hold part of the context window. While it doesn't speed up writing below 64k, the iPhone assists with writing beyond 64k, preventing the Mac from needing to drop to 4-bit context for 128k. The developer's custom kernels and DFlash2 speculative decoding also boosted token generation from 11.3 tok/s to 25 tok/s at 30k context.
This report details a novel method to offload GPU tasks to an iPhone, achieving 29-44% faster prefilling for Qwen 3.8 27B, unlike typical MacBook setups.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 3, 2026, 02:00 UTC
- Ingested
- Oct 3, 2026, 02:00
- Source type
- Dev community
Full text isn't available here.
Read at source →