Skip to content
RCreddit.com·

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

AI summary

The new Nex-N2.5-mini-MLX-4bit model, deployable on consumer hardware, has been benchmarked on an Apple M5 Max, achieving 133.6 tok/s. Using recommended settings like temperature: 0.7, top_p: 0.95, top_k: 40, and reasoning_effort: high, the model demonstrated good generation speed, prompt processing, and quality, with an acceptable memory footprint. Benchmarks are available on llm-bench.io, and the quant used is from Hugging Face.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 12, 2026, 19:00 UTC

IngestedOffset at this time: UTC+0Sep 13, 2026, 15:01 UTC

Published
Sep 12, 2026, 19:00
Ingested
Sep 13, 2026, 15:01
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini

I went with the recommended settings for the best generation quality and ran a few benchmarks:

I must say, the outcome is not bad at all - really good generation speed and prompt processing, okay memory footprint and good quality across the board. Will for sure give it a try to fuel my agents and might also try to do some coding with it. All benchmarks run I did you can find here: https://llm-bench.io/models/nex-n2-5-mini-mlx-4bit

Source·reddit.com