Skip to content
RCreddit.com·
Not on the current live radar

Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

AI summary

An experiment successfully transferred Qwen3.8 Flash-Next’s pretrained PLE n-gram memory into a smaller Qwen3.5-0.8B model, resulting in a 5.05% lower validation perplexity. The setup involved freezing both the Qwen3.5-0.8B backbone and the 51B-parameter PLE memory. A small R=1 reader was trained at decoder layers 3 and 9, with a token-dependent linear gate controlling the injection. The author, who developed Qwengram, designed the experiments and is responsible for the conclusions.

Why this one

This report details the first successful transfer of n-gram memory from a larger Qwen3.8 Flash-Next model to a smaller Qwen3.5-0.8B, unlike prior attempts that focused on other architectural components.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 25, 2026, 16:02 UTC

Ingested
Sep 25, 2026, 16:02
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com