Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
2025/05/27 by Boyang Wang, Xuweiyi Chen, Wang, Boyang +5 · 4 citations
Computer Science · Engineering · #Advanced Vision and Imaging #Image Processing Techniques and Applications #Computer Graphics and Visualization Techniques
paper · pdf · doi:10.48550/arxiv.2505.21491
Abstract
Controllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cinematic technique known as Frame In and Frame Out. Specifically, starting from image-to-video generation, users can control the objects in the image to naturally leave the scene or provide breaking new identity references to enter the scene, guided by a user-specified motion trajectory. To support this task, we introduce a new dataset that is curated semi-automatically, an efficient identity-preserving motion-controllable video Diffusion Transformer architecture, and a comprehensive evaluation protocol targeting this task. Our evaluation shows that our proposed approach significantly outperforms existing baselines.
Citations
- FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
- SkyReels-A2: Compose Anything in Video Diffusion Transformers
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
- Wan: Open and Advanced Large-Scale Video Generative Models
- FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
- Concat-ID: Towards Universal Identity-Preserving Video Synthesis
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- ObjectMover: Generative Object Movement with Video Prior
- EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
- VACE: All-in-One Video Creation and Editing
- Get In Video: Add Anything You Want to the Video
- VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
- Phantom: Subject-consistent video generation via cross-modal alignment
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
- History-Guided Video Diffusion
- MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation
- Continuous 3D Perception Model with Persistent State
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
- TransPixeler: Advancing Text-to-Video Generation with Transparency
- Cosmos World Foundation Model Platform for Physical AI
- VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
- LTX-Video: Realtime Video Latent Diffusion
- LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis
- Qwen2.5 Technical Report
- MotionBridge: Dynamic Video Inbetweening with Flexible Controls
- ObjCtrl-2.5D: Training-free Object Control with Camera Poses
- 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
- MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Identity-Preserving Text-to-Video Generation by Frequency Decomposition
- OminiControl: Minimal and Universal Control for Diffusion Transformer
- Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
- SAM 2: Segment Anything in Images and Videos
- Tora: Trajectory-oriented Diffusion Transformer for Video Generation
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
- Image Conductor: Precision Control for Interactive Video Synthesis
- One-Step Effective Diffusion Network for Real-World Image Super-Resolution
- ToonCrafter: Generative Cartoon Interpolation
- ReVideo: Remake a Video with Motion and Content Control
- Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- Video Interpolation with Diffusion Models
- StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
- AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
- Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation
- BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
- APISR: Anime Production Inspired Real-World Anime Super-Resolution
- Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
- InstantID: Zero-shot Identity-Preserving Generation in Seconds
- PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
- MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
- Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- ProPainter: Improving Propagation and Transformer for Video Inpainting
- DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
- Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
- DINOv2: Learning Robust Visual Features without Supervision
- DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion
- Adding Conditional Control to Text-to-Image Diffusion Models
- Adding Conditional Control to Text-to-Image Diffusion Models
- Scalable Diffusion Models with Transformers
- OneFormer: One Transformer to Rule Universal Image Segmentation
- Classifier-Free Diffusion Guidance
- Exploring CLIP for Assessing the Look and Feel of Images
- 8-bit Optimizers via Block-wise Quantization
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- TransNet V2: An effective deep network architecture for fast shot transition detection
- Character Region Awareness for Text Detection
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Panoptic Segmentation
- NIMA: Neural Image Assessment
- Microsoft COCO: Common Objects in Context
Cited by
Related