Engram gone wild! 2b model update...
An update on a 2.6b model with a 4.3b ENGRAM table has been released, detailing its architecture and training methodology. The model uses SWA layering 3:1 and incorporates a context-aware Engram table after the first four layers, utilizing a small set of attention heads for data injection. Attention Based Residuals (Moonshot) allow subsequent blocks to interact with the Engram table. The LM Head and Embed layers are frozen, down-projected via SVD from OLMo 3's model, saving significant training time. The data is stored in a ~KD format, using 32 soft targets from a Wikipedia corpus teacher model.
This update details a 2.6b model with an unusually large 4.3b Engram table, a mismatch of compute vs. table size that the developer notes is unlike any other model they know of.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 23, 2026, 09:01 UTC
- Ingested
- Sep 23, 2026, 09:01
- Source type
- Dev community
Full text isn't available here.
Read at source →