Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io
The new Nex-N2.5-mini-MLX-4bit model, deployable on consumer hardware, has been benchmarked on an Apple M5 Max, achieving 133.6 tok/s. Using recommended settings like temperature: 0.7, top_p: 0.95, top_k: 40, and reasoning_effort: high, the model demonstrated good generation speed, prompt processing, and quality, with an acceptable memory footprint. Benchmarks are available on llm-bench.io, and the quant used is from Hugging Face.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 12, 2026, 19:00 UTC
IngestedOffset at this time: UTC+0Sep 13, 2026, 15:01 UTC
- Published
- Sep 12, 2026, 19:00
- Ingested
- Sep 13, 2026, 15:01
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini
I went with the recommended settings for the best generation quality and ran a few benchmarks:
I must say, the outcome is not bad at all - really good generation speed and prompt processing, okay memory footprint and good quality across the board. Will for sure give it a try to fuel my agents and might also try to do some coding with it. All benchmarks run I did you can find here: https://llm-bench.io/models/nex-n2-5-mini-mlx-4bit