Skip to content
RCreddit.com·

I wrote a free, open-source book on making ML models actually fast, from silicon to agents [P]

AI summary

A new free, open-source book titled "How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents" has been released. The author spent several months writing this resource, aiming to provide insights into ML performance engineering that they wished they had when starting out. The book covers methods for making ML models fast, from silicon to agents, and the author appreciates a star if readers find it useful.

Why this one

This book offers a unique, comprehensive systems view of ML performance, unlike many resources that focus on isolated aspects of model optimization.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 29, 2026, 10:35 UTC

IngestedOffset at this time: UTC+0Sep 29, 2026, 15:00 UTC

Published
Sep 29, 2026, 10:35
Ingested
Sep 29, 2026, 15:00
Source type
Dev community
Tier
Community
Source status
Sync delayed

Tier is a per-source editorial setting, not a per-item score.

I’ve spent the last few months writing something I wish I had when I started working on ML performance engineering.

It’s called How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents.

The basic idea is that reducing FLOPs doesn’t necessarily make a model faster. Before optimising anything, you need to understand what the system is actually bounded by.

The book starts with roofline analysis and hardware, then works its way up through kernels, compilers, quantisation, pruning, vision, on-device LLMs, robotics, profiling, serving and finally agents.

The goal is to build the intuition to look at a model and a piece of hardware and reason about:

- How fast can this possibly run?

- Am I compute, bandwidth, memory or system bound?

- Which optimisation will actually move that limit?

- Is quantisation, pruning or kernel optimisation even worth doing here?

- What happens when the same thinking is applied to serving and agent systems?

The whole thing is free and open source:

https://github.com/usamahz/make-your-model-fast

Would genuinely appreciate feedback or contributions from people working on ML systems, inference, compilers, edge AI or performance engineering.

And if you find it useful, a ⭐ would be appreciated!

Source·reddit.com