LTX-Video: Realtime Video Latent Diffusion
2024/12/30 by Yoav HaCohen, HaCohen, Yoav, Nisan Chiprut +29 · 2 voices · 200 citations
Computer Science · #Advanced Data Compression Techniques #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #cs.CV
paper · pdf · doi:10.48550/arxiv.2501.00103
openalex publication_date 2024/12/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce LTX-Video, a transformer-based latent diffusion model that adopts a holistic approach to video generation by seamlessly integrating the responsibilities of the Video-VAE and the denoising transformer. Unlike existing methods, which treat these components as independent, LTX-Video aims to optimize their interaction for improved efficiency and quality. At its core is a carefully designed Video-VAE that achieves a high compression ratio of 1:192, with spatiotemporal downscaling of 32 x 32 x 8 pixels per token, enabled by relocating the patchifying operation from the transformer's input to the VAE's input. Operating in this highly compressed latent space enables the transformer to efficiently perform full spatiotemporal self-attention, which is essential for generating high-resolution videos with temporal consistency. However, the high compression inherently limits the representation of fine details. To address this, our VAE decoder is tasked with both latent-to-pixel conversion and the final denoising step, producing the clean result directly in pixel space. This approach preserves the ability to generate fine details without incurring the runtime cost of a separate upsampling module. Our model supports diverse use cases, including text-to-video and image-to-video generation, with both capabilities trained simultaneously. It achieves faster-than-real-time generation, producing 5 seconds of 24 fps video at 768x512 resolution in just 2 seconds on an Nvidia H100 GPU, outperforming all existing models of similar scale. The source code and pre-trained models are publicly available, setting a new benchmark for accessible and scalable video generation.
Cited by
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
- Act2Goal: From World Model To General Goal-conditioned Policy
- Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
- MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
- MobileWan: Closing the Quality Gap for Mobile Video Diffusion
- Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
- CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion
- AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- SVBench: Evaluation of Video Generation Models on Social Reasoning
- HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
- DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
- WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
- Over++: Generative Video Compositing for Layer Interaction Effects
- Region-Constraint In-Context Generation for Instructional Video Editing
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Spatia: Video Generation with Updatable Spatial Memory
- MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- What Happens Next? Next Scene Prediction with a Unified Video Model
- Animus3D: Text-driven 3D Animation via Motion Score Distillation
- Generative Spatiotemporal Data Augmentation
- Endless World: Real-Time 3D-Aware Long Video Generation
- CineLOG: A Training Free Approach for Cinematic Long Video Generation
- V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
- Flowception: Temporally Expansive Flow Matching for Video Generation
- Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
- REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- Self-Evolving 3D Scene Generation from a Single Image
- EgoX: Egocentric Video Generation from a Single Exocentric Video
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
- VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- Scaling Zero-Shot Reference-to-Video Generation
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
- Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
- MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- Spatiotemporal Pyramid Flow Matching for Climate Emulation
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
- IC-World: In-Context Generation for Shared World Modeling
- What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
- DisMo: Disentangled Motion Representations for Open-World Motion Transfer
- One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
- Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns
- Reinforcing Action Policies by Prophesying
- Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
- Loomis Painter: Reconstructing the Painting Process
- Decoupling Complexity from Scale in Latent Diffusion Model
- Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine
- Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
- Neodragon: Mobile Video Generation using Diffusion Transformer
- Towards One-step Causal Video Generation via Adversarial Self-Distillation
- VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
- G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
- Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
- Rethinking Visual Intelligence: Insights from Video Pretraining
- World Simulation with Video Foundation Models for Physical AI
- BachVid: Training-Free Video Generation with Consistent Background and Character
- Epipolar Geometry Improves Video Generation Models
- PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation
- World-in-World: World Models in a Closed-Loop World
- Latent Diffusion Model without Variational Autoencoder
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
- Asymmetric Flow Models
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
- PickStyle: Video-to-Video Style Transfer with Context-Style Adapters
- DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
- Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report
- Learning to Generate Rigid Body Interactions with Video Diffusion Models
- EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
- InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
- DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
- UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- Autoregressive Video Generation beyond Next Frames Prediction
- VideoScore2: Think before You Score in Generative Video Evaluation
- LongLive: Real-time Interactive Long Video Generation
- Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
- Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
- What Happens Next? Anticipating Future Motion by Generating Point Trajectories
- ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
- MAD: Motion Appearance Decoupling for efficient Driving World Models
- LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning
- DiCache: Let Diffusion Model Determine Its Own Cache
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
- PlayerOne: Egocentric World Simulator
- CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
- Mixture of Contexts for Long Video Generation
- Phased One-Step Adversarial Equilibrium for Video Diffusion Models
- Matrix-game 2.0: An open-source real-time and streaming interactive world model
- Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
- DreamVE: Unified Instruction-based Image and Video Editing
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
- Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
- Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
- "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- LoViC: Efficient Long Video Generation with Context Compression
- T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
- Captain Cinema: Towards Short Movie Generation
- Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
- Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
- Retrieval-Driven Training-Free AI-Generated Video Attribution
- MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
- MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
- Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
- DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
- LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
- FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
- TurboVSR: Fantastic Video Upscalers and Where to Find Them
- JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
- Listener-Rewarded Thinking in VLMs for Image Preferences
- Radial Attention: O(nlog n) Sparse Attention with Energy Decay for Long Video Generation
- GenHSI: Controllable Generation of Human-Scene Interaction Videos
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
- Emergent Temporal Correspondences from Video Diffusion Transformers
- FramePrompt: In-context Controllable Animation with Zero Structural Changes
- From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
- EraserDiT: Fast Video Inpainting with Diffusion Transformer Model
- Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
- Video World Models with Long-term Spatial Memory
- Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
- Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
- Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
- Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis
- MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
- LatentMove: Towards Complex Human Movement Video Generation
- Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
- Long-Context State-Space Video World Models
- AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
- Model Already Knows the Best Noise: Bayesian Active Noise Selection via Attention in Video Diffusion Model
- Vid2World: Crafting Video Diffusion Models to Interactive World Models
- BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation
- VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
- T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
- Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
- Infinite Worlds with Versatile Interactions
- FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion
- T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
- Latent Spatial Memory for Video World Models
- Helios: Real Real-Time Long Video Generation Model
- OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
- Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
- VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
- Image Generation with a Sphere Encoder
- SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
- Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
- Learning Long-term Motion Embeddings for Efficient Kinematics Generation
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- Latent-Compressed Variational Autoencoder for Video Diffusion Models
- VideoMaMa: Mask-Guided Video Matting via Generative Prior
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
- Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
- Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
- Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute
- Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform
- Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
- VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
- In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
- Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
- H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
- AB-Cache: Training-Free Acceleration of Diffusion Models via Adams-Bashforth Cached Feature Reuse
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
Discussions
Related