2025/06/14 by Saemee Choi, Choi, Saemee, Sohyun Jeong +6
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.2506.12520
openalex publication_date 2025/06/14 · openalex created_date 2025/10/13 · openalex updated_date 2026/07/28
We propose VINO, the first zero-shot, training-free video editing method conditioned on both image and text. Our approach introduces ρ-start sampling and dilated dual masking to construct structured noise maps that enable coherent and accurate edits. To further enhance visual fidelity, we present zero image guidance, a controllable negative prompt strategy. Extensive experiments demonstrate that VINO faithfully incorporates the reference image into video edits, achieving strong performance compared to state-of-the-art baselines, all without any test-time or instance-specific training.