Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io
The new Nex-N2.5-mini-MLX-4bit model, deployable on consumer hardware, has been benchmarked on an Apple M5 Max, achieving 133.6 tok/s. Using recommended settings like temperature: 0.7, top_p: 0.95, top_k: 40, and reasoning_effort: high, the model demonstrated good generation speed, prompt processing, and quality, with an acceptable memory footprint. Benchmarks are available on llm-bench.io, and the quant used is from Hugging Face.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月12日 19:00 UTC
收录当时偏移:UTC+02026年9月13日 15:01 UTC
- 发布
- 2026年9月12日 19:00
- 收录
- 2026年9月13日 15:01
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini
I went with the recommended settings for the best generation quality and ran a few benchmarks:
I must say, the outcome is not bad at all - really good generation speed and prompt processing, okay memory footprint and good quality across the board. Will for sure give it a try to fuel my agents and might also try to do some coding with it. All benchmarks run I did you can find here: https://llm-bench.io/models/nex-n2-5-mini-mlx-4bit