I wrote a free, open-source book on making ML models actually fast, from silicon to agents [P]
A new free, open-source book titled "How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents" has been released. The author spent several months writing this resource, aiming to provide insights into ML performance engineering that they wished they had when starting out. The book covers methods for making ML models fast, from silicon to agents, and the author appreciates a star if readers find it useful.
This book offers a unique, comprehensive systems view of ML performance, unlike many resources that focus on isolated aspects of model optimization.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月29日 10:35 UTC
收录当时偏移:UTC+02026年9月29日 15:00 UTC
- 发布
- 2026年9月29日 10:35
- 收录
- 2026年9月29日 15:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
I’ve spent the last few months writing something I wish I had when I started working on ML performance engineering.
It’s called How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents.
The basic idea is that reducing FLOPs doesn’t necessarily make a model faster. Before optimising anything, you need to understand what the system is actually bounded by.
The book starts with roofline analysis and hardware, then works its way up through kernels, compilers, quantisation, pruning, vision, on-device LLMs, robotics, profiling, serving and finally agents.
The goal is to build the intuition to look at a model and a piece of hardware and reason about:
- How fast can this possibly run?
- Am I compute, bandwidth, memory or system bound?
- Which optimisation will actually move that limit?
- Is quantisation, pruning or kernel optimisation even worth doing here?
- What happens when the same thinking is applied to serving and agent systems?
The whole thing is free and open source:
https://github.com/usamahz/make-your-model-fast
Would genuinely appreciate feedback or contributions from people working on ML systems, inference, compilers, edge AI or performance engineering.
And if you find it useful, a ⭐ would be appreciated!