UniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement
UniEvo-VL is a self-evolving framework for multimodal models that uses self-correction feedback during test-time compute. It allows a single multimodal model to act as both teacher and student, with the student seeing the vanilla question and the teacher conditioning on privileged critiques. Training minimizes per-state divergence between their denoising diffusion distributions over the student's sampling trajectories. Experiments show UniEvo-VL improves image generation capabilities, with performance gains on GenEval (0.747 to 0.808) and GenEval2 Soft-TIFA (32.97 to 35.53) using Qwen-image-2512.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 7, 2026, 01:00 UTC
- Ingested
- Oct 7, 2026, 01:00
- Source type
- Research
Full text isn't available here.
Read at source →