RCreddit.com·
暂不在当前实时榜单
Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact
Model grafting transforms existing models like Qwen3.5-4B into causal encoder-decoders, a technique distinct from training such architectures from scratch, as seen with DeepSeek-V4.1-Flash. This method involves cutting the model at a certain depth, using lower layers for prompt reading, and upper layers for the encoder's residual stream via identity-init adapters, then healing with self-distillation. Two Qwen3.5-4B graft variants, graft8 and graft16, were created. Graft8 achieves a ~3.7x speedup at 128K prompt with some accuracy loss, while graft16 offers a 2.0x speedup with minimal accuracy loss.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月22日 19:01 UTC
- 收录
- 2026年9月22日 19:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →