How I got my Mac to read my Claude Code chats at night and extend my token usage by 1/3rd
Heat trend
The percentage is based on available heat signal, not comment count or independent people.
Claude model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.
A developer found a way to extend their token usage by one-third by having their Mac locally read Claude Code chat histories. Previously, a significant portion of their token usage was spent rereading unchanged files and old chat logs. Now, the Mac processes chats as they happen, summarizing key information for new conversations, effectively reducing the need for the expensive model to reread everything. This method allows them to maintain their Max plan while making token usage more efficient.
I kept hitting my weekly limit two days before my reset. So I dug into where the week was actually going. About a third of my usage was re-reads: the same files getting read again even though they hadn't changed, and old chats getting pulled into new ones over and over.
Apple Intelligence might not be the fastest AI, but it's great at reading through my chats at night for free. You don't need the smartest model to read a chat log. Reading is easy work.
So my setup now: my Mac reads each chat as it grows, on my own machine, zero tokens, even while I sleep. When a chat is finished, my Mac writes the handoff: what happened, what was decided, what's next, which files were touched. The next morning, my new chat starts already knowing all of that, instead of the expensive model re-reading everything at full price. And I have it set up so anything my Mac can't read gets pushed to Haiku. So far I haven't needed it, my Mac has handled all of it.
I kept my same Max plan but I'm making the tokens go about 1/3rd further.