Skip to content
HNHacker News·
Archived topic · source no longer tracked

Show HN: Getting GLM 5.2 running on my slow computer

AI summary

A new project, Colibrì, enables running the 744B-parameter GLM-5.2 Mixture-of-Experts model on consumer machines with 25 GB RAM. It achieves this by streaming experts from disk, keeping only the dense part of the model (9.9 GB) resident in RAM. The pure C engine, with zero dependencies, utilizes techniques like MLA attention, DeepSeek-V3-style routing, and native MTP speculative decoding for efficient operation, despite cold starts being slow due to disk reads.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Jul 10, 2026, 05:00 UTC

Ingested
Jul 10, 2026, 05:00
Source type
Unclassified