Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact
Model grafting transforms existing models like Qwen3.5-4B into causal encoder-decoders, a technique distinct from training such architectures from scratch, as seen with DeepSeek-V4.1-Flash. This method involves cutting the model at a certain depth, using lower layers for prompt reading, and upper layers for the encoder's residual stream via identity-init adapters, then healing with self-distillation. Two Qwen3.5-4B graft variants, graft8 and graft16, were created. Graft8 achieves a ~3.7x speedup at 128K prompt with some accuracy loss, while graft16 offers a 2.0x speedup with minimal accuracy loss.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 22, 2026, 19:01 UTC
- Ingested
- Sep 22, 2026, 19:01
- Source type
- Dev community
Full text isn't available here.
Read at source →