Imagen Video: High Definition Video Generation with Diffusion Models
2022/10/05 by Jonathan Ho, William Chan, Ho, Jonathan +19 · 145 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Computer Graphics and Visualization Techniques #Advanced Image Processing Techniques
paper · pdf · doi:10.48550/arxiv.2210.02303
Abstract
We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models. Given a text prompt, Imagen Video generates high definition videos using a base video generation model and a sequence of interleaved spatial and temporal video super-resolution models. We describe how we scale up the system as a high definition text-to-video model including design decisions such as the choice of fully-convolutional temporal and spatial super-resolution models at certain resolutions, and the choice of the v-parameterization of diffusion models. In addition, we confirm and transfer findings from previous work on diffusion-based image generation to the video generation setting. Finally, we apply progressive distillation to our video models with classifier-free guidance for fast, high quality sampling. We find Imagen Video not only capable of generating videos of high fidelity, but also having a high degree of controllability and world knowledge, including the ability to generate diverse videos and text animations in various artistic styles and with 3D object understanding. See https://imagen.research.google/video/ for samples.
Cited by
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
- PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image Generation
- LangPrecip: Language-Aware Multimodal Precipitation Nowcasting
- Generalized Fine-Tuning of Diffusion Models via Stochastic Control and FBSDEs
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
- LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs
- SemanticGen: Video Generation in Semantic Space
- WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
- ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
- In-Context Audio Control of Video Diffusion Transformers
- Preserving Spectral Structure and Statistics in Diffusion Models
- Region-Constraint In-Context Generation for Instructional Video Editing
- Large Video Planner Enables Generalizable Robot Control
- Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
- DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
- AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
- Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces
- Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis
- Splatent: Splatting Diffusion Latents for Novel View Synthesis
- WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
- SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
- Glance: Accelerating Diffusion Models with 1 Sample
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- Spatiotemporal Pyramid Flow Matching for Climate Emulation
- InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
- What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
- DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
- Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
- MotionV2V: Editing Motion in a Video
- STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
- Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- Learning Plug-and-play Memory for Guiding Video Diffusion Models
- Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- Point-to-Point: Sparse Motion Guidance for Controllable Video Editing
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
- Planning with Sketch-Guided Verification for Physics-Aware Video Generation
- MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine
- VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- Neodragon: Mobile Video Generation using Diffusion Transformer
- RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation
- PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
- Dexterous Robotic Piano Playing at Scale
- Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides
- A Sensing Whole Brain Zebrafish Foundation Model for Neuron Dynamics and Behavior
- A Step Toward World Models: A Survey on Robotic Manipulation
- AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency
- TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models
- Walk through Paintings: Egocentric World Models from Internet Priors
- See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
- M3T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark
- Improved Training Technique for Shortcut Models
- HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- BadGraph: A Backdoor Attack Against Latent Diffusion Model for Text-Guided Graph Generation
- UltraGen: High-Resolution Video Generation with Hierarchical Attention
- LAND: Lung and Nodule Diffusion for 3D Chest CT Synthesis with Anatomical Guidance
- Attention Is All You Need for KV Cache in Diffusion LLMs
- Inference-Time Search using Side Information for Diffusion-based Image Reconstruction
- Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation
- VIDMP3: Video Editing by Representing Motion with Pose and Position Priors
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- DiffStyleTS: Diffusion Model for Style Transfer in Time Series
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- MultiCOIN: Multi-Modal COntrollable Video INbetweening
- SummDiff: Generative Modeling of Video Summarization with Diffusion
- VideoVerse: How Far is Your T2V Generator from a World Model?
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- Real-Time Motion-Controllable Autoregressive Video Diffusion
- Ctrl-VI: Controllable Video Synthesis via Variational Inference
- Trajectory Conditioned Cross-embodiment Skill Transfer
- Reinforcing Diffusion Models by Direct Group Preference Optimization
- Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
- Diffusion Models and the Manifold Hypothesis: Log-Domain Smoothing is Geometry Adaptive
- Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
- VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
- Character Mixing for Video Generation
- Diffusion2: Turning 3D Environments into Radio Frequency Heatmaps
- IMAGEdit: Let Any Subject Transform
- Efficient Probabilistic Tensor Networks
- Code2Video: A Code-centric Paradigm for Educational Video Generation
- EVODiff: Entropy-aware Variance Optimized Diffusion Inference
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
- UniVid: The Open-Source Unified Video Model
- Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- Autoregressive Video Generation beyond Next Frames Prediction
- LLM/Agent-as-Data-Analyst: A Survey
- SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
- Diff-3DCap: Shape Captioning with Diffusion Models
- RestoRect: Degraded Image Restoration via Latent Rectified Flow & Feature Distillation
- Score-based Idempotent Distillation of Diffusion Models
- VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- ControlEchoSynth: Boosting Ejection Fraction Estimation Models via Controlled Video Diffusion
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
- Every Camera Effect, Every Time, All at Once: 4D Gaussian Ray Tracing for Physics-based Camera Effect Data Generation
- T2Bs: Text-to-Character Blendshapes via Video Generation
- Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
- ANYPORTAL: Zero-Shot Consistent Video Background Replacement
- Zero-shot 3D-Aware Trajectory-Guided image-to-video generation via Test-Time Training
- DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
- STADI: Fine-Grained Step-Patch Diffusion Parallelism for Heterogeneous GPUs
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
- Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
- Complete Gaussian Splats from a Single Image with Denoising Diffusion Models
- PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
- Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
- InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
- 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
- CTFlow: Video-Inspired Latent Flow Matching for 3D CT Synthesis
- LIA-X: Interpretable Latent Portrait Animator
- Generation of Indian Sign Language Letters, Numbers, and Words
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- Learning an Implicit Physics Model for Image-based Fluid Simulation
- S2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
- DiTVR: Zero-Shot Diffusion Transformer for Video Restoration
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Versatile Transition Generation with Image-to-Video Diffusion
- Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- GVD: Guiding Video Diffusion Model for Scalable Video Distillation
- Low-Cost Test-Time Adaptation for Robust Video Editing
Related