Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity
An experiment successfully transferred Qwen3.8 Flash-Next’s pretrained PLE n-gram memory into a smaller Qwen3.5-0.8B model, resulting in a 5.05% lower validation perplexity. The setup involved freezing both the Qwen3.5-0.8B backbone and the 51B-parameter PLE memory. A small R=1 reader was trained at decoder layers 3 and 9, with a token-dependent linear gate controlling the injection. The author, who developed Qwengram, designed the experiments and is responsible for the conclusions.
This report details the first successful transfer of n-gram memory from a larger Qwen3.8 Flash-Next model to a smaller Qwen3.5-0.8B, unlike prior attempts that focused on other architectural components.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月25日 16:02 UTC
- 收录
- 2026年9月25日 16:02
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →