跳到正文
RCreddit.com·
暂不在当前实时榜单

Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

AI 摘要

An experiment successfully transferred Qwen3.8 Flash-Next’s pretrained PLE n-gram memory into a smaller Qwen3.5-0.8B model, resulting in a 5.05% lower validation perplexity. The setup involved freezing both the Qwen3.5-0.8B backbone and the 51B-parameter PLE memory. A small R=1 reader was trained at decoder layers 3 and 9, with a token-dependent linear gate controlling the injection. The author, who developed Qwengram, designed the experiments and is responsible for the conclusions.

为什么是这条

This report details the first successful transfer of n-gram memory from a larger Qwen3.8 Flash-Next model to a smaller Qwen3.5-0.8B, unlike prior attempts that focused on other architectural components.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月25日 16:02 UTC

收录
2026年9月25日 16:02
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com