Back
RCreddit.com
15
·18 hr ago·Dev community · RSS

Mac Heads: Is there any point to MLX in September 2026?

View original
Model release

Heat trend

New
Latest 24h versus previous 24h · 7-day curve

The percentage is based on available heat signal, not comment count or independent people.

AI summary

A Reddit user is questioning the future relevance of MLX for Mac users, specifically in September 2026, when running models like Qwen3.8 27b on Apple M5 Pro chips. They are seeking to understand if their observed performance of ~19t/s generation and 300t/s+ prefill with Q8/8-bit quantization is optimal, or if there are methods to achieve better results or utilize mainstream MLX MTP effectively. The user also wonders if the specific model, Qwen3.8 27b, is a limiting factor.

This may be somewhat specific to Qwen3.8 27b and the Apple M5 series, perhaps, but enough of us are running this combo that it's worth tossing out there.

GGUF models with MTP have been the fastest way to go for some time for token generation, except possibly for a few tweaked MTPLX models running on alpha-stage MLX forks. Prefill, however, was still much faster for M5s under MLX. This has caused me to switch models depending on the expected generation/prefill mix, which is annoying.

While I wasn't looking, it appears that llama.cpp for Metal must have added support for M5 matmul/"neural accelerators" because prefill performance with e.g. Unsloth's Q_8 GGUF is now at least as good (~300-350t/s) as anything I have seen with MLX models--even in oMLX. This was a pleasant surprise! Now I can't think of a reason to use MLX models at all.

Am I missing something? Are my observation bogus? Could I do better than ~19t/s generation and 300t/s+ prefill on a M5 Pro with Qwen3.8 27b in the Q8/8-bit range? Is there a secret handshake to get mainstream MLX MTP working?

Or is this just because of the specific model in question?

Mac Heads: Is there any point to MLX in September 2026? · BuzzRadr