Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity
An experiment successfully transferred Qwen3.8 Flash-Next’s pretrained PLE n-gram memory into a smaller Qwen3.5-0.8B model, resulting in a 5.05% lower validation perplexity. The setup involved freezing both the Qwen3.5-0.8B backbone and the 51B-parameter PLE memory. A small R=1 reader was trained at decoder layers 3 and 9, with a token-dependent linear gate controlling the injection. The author, who developed Qwengram, designed the experiments and is responsible for the conclusions.
This report details the first successful transfer of n-gram memory from a larger Qwen3.8 Flash-Next model to a smaller Qwen3.5-0.8B, unlike prior attempts that focused on other architectural components.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 25, 2026, 16:02 UTC
- Ingested
- Sep 25, 2026, 16:02
- Source type
- Dev community
Full text isn't available here.
Read at source →