vix.ing · top · new · best · stats · spec

Video4Edit: Viewing Image Editing as a Degenerate Temporal Process

2025/11/22 by Li, Xiaofan, Sun, Yanpeng, Wu, Chenming +5
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Visual Attention and Saliency Detection

paper · doi:10.48550/arxiv.2511.18131

openalex publication_date 2025/11/22 · openalex created_date 2025/11/27 · openalex updated_date 2026/07/28

Abstract

We observe that recent advances in multimodal foundation models have propelled instruction-driven image generation and editing into a genuinely cross-modal, cooperative regime. Nevertheless, state-of-the-art editing pipelines remain costly: beyond training large diffusion/flow models, they require curating massive high-quality triplets of \instruction, source image, edited image\ to cover diverse user intents. Moreover, the fidelity of visual replacements hinges on how precisely the instruction references the target semantics. We revisit this challenge through the lens of temporal modeling: if video can be regarded as a full temporal process, then image editing can be seen as a degenerate temporal process. This perspective allows us to transfer single-frame evolution priors from video pre-training, enabling a highly data-efficient fine-tuning regime. Empirically, our approach matches the performance of leading open-source baselines while using only about one percent of the supervision demanded by mainstream editing models.

Citations

Related