Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
2025/04/11 by Jialu Li, Li, Jialu, Shoubin Yu +9 · 10 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation #Natural Language Processing Techniques #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2504.08641
openalex publication_date 2025/04/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recent advancements in text-to-video (T2V) diffusion models have significantly enhanced the visual quality of the generated videos. However, even recent T2V models find it challenging to follow text descriptions accurately, especially when the prompt requires accurate control of spatial layouts or object trajectories. A recent line of research uses layout guidance for T2V models that require fine-tuning or iterative manipulation of the attention map during inference time. This significantly increases the memory requirement, making it difficult to adopt a large T2V model as a backbone. To address this, we introduce Video-MSG, a training-free Guidance method for T2V generation based on Multimodal planning and Structured noise initialization. Video-MSG consists of three steps, where in the first two steps, Video-MSG creates Video Sketch, a fine-grained spatio-temporal plan for the final video, specifying background, foreground, and object trajectories, in the form of draft video frames. In the last step, Video-MSG guides a downstream T2V diffusion model with Video Sketch through noise inversion and denoising. Notably, Video-MSG does not need fine-tuning or attention manipulation with additional memory during inference time, making it easier to adopt large T2V models. Video-MSG demonstrates its effectiveness in enhancing text alignment with multiple T2V backbones (VideoCrafter2 and CogVideoX-5B) on popular T2V generation benchmarks (T2VCompBench and VBench). We provide comprehensive ablation studies about noise inversion ratio, different background generators, background object detection, and foreground object segmentation.
Citations
- Wan: Open and Advanced Large-Scale Video Generative Models
- MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent
- VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
- Cosmos World Foundation Model Platform for Physical AI
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Movie Gen: A Cast of Media Foundation Models
- Emu3: Next-Token Prediction is All You Need
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- TrackGo: A Flexible and Efficient Method for Controllable Video Generation
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Tora: Trajectory-oriented Diffusion Transformer for Video Generation
- T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
- Image Conductor: Precision Control for Interactive Video Synthesis
- Compositional Video Generation as Flow Equalization
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- Genie: Generative Interactive Environments
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
- Latte: Latent Diffusion Transformer for Video Generation
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- VideoPoet: A Large Language Model for Zero-Shot Video Generation
- MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- LLM-grounded Video Diffusion Models
- Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
- VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
- DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory
- ModelScope Text-to-Video Technical Report
- Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Recognize Anything: A Strong Image Tagging Model
- Micrograph segmentations for DDEVD
- Segment Anything
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
- Denoising Diffusion Implicit Models
- Denoising Diffusion Probabilistic Models
- DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
Cited by
Related