跳到正文
HNHacker News·
暂不在当前实时榜单

UniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement

AI 摘要

UniEvo-VL is a self-evolving framework for multimodal models that uses self-correction feedback during test-time compute. It allows a single multimodal model to act as both teacher and student, with the student seeing the vanilla question and the teacher conditioning on privileged critiques. Training minimizes per-state divergence between their denoising diffusion distributions over the student's sampling trajectories. Experiments show UniEvo-VL improves image generation capabilities, with performance gains on GenEval (0.747 to 0.808) and GenEval2 Soft-TIFA (32.97 to 35.53) using Qwen-image-2512.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月7日 01:00 UTC

收录
2026年10月7日 01:00
来源类型
研究

本站未收录正文。

前往源站阅读 →
来源·Hacker News·arxiv.org