Engram gone wild! 2b model update...
An update on a 2.6b model with a 4.3b ENGRAM table has been released, detailing its architecture and training methodology. The model uses SWA layering 3:1 and incorporates a context-aware Engram table after the first four layers, utilizing a small set of attention heads for data injection. Attention Based Residuals (Moonshot) allow subsequent blocks to interact with the Engram table. The LM Head and Embed layers are frozen, down-projected via SVD from OLMo 3's model, saving significant training time. The data is stored in a ~KD format, using 32 soft targets from a Wikipedia corpus teacher model.
This update details a 2.6b model with an unusually large 4.3b Engram table, a mismatch of compute vs. table size that the developer notes is unlike any other model they know of.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月23日 09:01 UTC
- 收录
- 2026年9月23日 09:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →