跳到正文
RCreddit.com·
暂不在当前实时榜单

Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact

AI 摘要

Model grafting transforms existing models like Qwen3.5-4B into causal encoder-decoders, a technique distinct from training such architectures from scratch, as seen with DeepSeek-V4.1-Flash. This method involves cutting the model at a certain depth, using lower layers for prompt reading, and upper layers for the encoder's residual stream via identity-init adapters, then healing with self-distillation. Two Qwen3.5-4B graft variants, graft8 and graft16, were created. Graft8 achieves a ~3.7x speedup at 128K prompt with some accuracy loss, while graft16 offers a 2.0x speedup with minimal accuracy loss.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月22日 19:01 UTC

收录
2026年9月22日 19:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com