Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
2024/02/27 by Yixin Liu, Kai Zhang, Liu, Yixin +21 · 3 voices · 105 citations
Earth and Planetary Sciences · Engineering · #3D Surveying and Cultural Heritage #Satellite Image Processing and Photogrammetry
paper · pdf · doi:10.48550/arxiv.2402.17177
Abstract
Sora is a text-to-video generative AI model, released by OpenAI in February 2024. The model is trained to generate videos of realistic or imaginative scenes from text instructions and show potential in simulating the physical world. Based on public technical reports and reverse engineering, this paper presents a comprehensive review of the model's background, related technologies, applications, remaining challenges, and future directions of text-to-video AI models. We first trace Sora's development and investigate the underlying technologies used to build this "world simulator". Then, we describe in detail the applications and potential impact of Sora in multiple industries ranging from film-making and education to marketing. We discuss the main challenges and limitations that need to be addressed to widely deploy Sora, such as ensuring safe and unbiased video generation. Lastly, we discuss the future development of Sora and video generation models in general, and how advancements in the field could enable new ways of human-AI interaction, boosting productivity and creativity of video generation.
Cited by
- KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
- GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
- SPEED: One-Step Pixel Diffusion for High-quality Video Frame Interpolation
- In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
- Flow Matching in Feature Space for Stochastic World Modeling
- Act2Goal: From World Model To General Goal-conditioned Policy
- Large Vision Model-Enhanced Digital Twin with Deep Reinforcement Learning for User Association and Load Balancing in Dynamic Wireless Networks
- Latent Space Probing for Adult Content Detection in Video Generative Models
- High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis
- Generating the Past, Present and Future from a Motion-Blurred Image
- STORM: Search-Guided Generative World Models for Robotic Manipulation
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- Factorized Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
- Toward Agentic Environments: GenAI and the Convergence of AI, Sustainability, and Human-Centric Spaces
- PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
- SneakPeek: Future-Guided Instructional Streaming Video Generation
- JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
- Self-Evolving 3D Scene Generation from a Single Image
- MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
- User Negotiations of Authenticity, Ownership, and Governance on AI-Generated Video Platforms: Evidence from Sora
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
- Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
- MindFuse: Towards GenAI Explainability in Marketing Strategy Co-Creation
- Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- MagicWorld: Interactive Geometry-driven Video World Exploration
- ObjectAlign: Neuro-Symbolic Object Consistency Verification and Correction
- Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
- Flow and Depth Assisted Video Prediction with Latent Transformer
- TS-PEFT: Unveiling Token-Level Redundancy in Parameter-Efficient Fine-Tuning
- Exact Stochastic Differential Equations for Quantum Reverse Diffusion
- First Frame Is the Place to Go for Video Content Customization
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
- From Events to Clarity: The Event-Guided Diffusion Framework for Dehazing
- GenAI vs. Human Creators: Procurement Mechanism Design in Two-/Three-Layer Markets
- Enhancing Diffusion Model Guidance through Calibration and Regularization
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
- Object-Aware 4D Human Motion Generation
- TridentServe: A Stage-level Serving System for Diffusion Pipelines
- NeurIPT: Foundation Model for Neural Interfaces
- MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
- Semantic Communications with World Models
- BachVid: Training-Free Video Generation with Consistent Background and Character
- VISTA: A Test-Time Self-Improving Video Generation Agent
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- From Mannequin to Human: A Pose-Aware and Identity-Preserving Video Generation Framework for Lifelike Clothing Display
- Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
- Inferring Dynamic Physical Properties from Video Foundation Models
- MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
- Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- From Noisy to Native: LLM-driven Graph Restoration for Test-Time Graph Domain Adaptation
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Code2Video: A Code-centric Paradigm for Educational Video Generation
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
- UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
- A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
- Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
- Drag4D: Align Your Motion with Text-Driven 3D Scene Generation
- MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
- MolMark: Safeguarding Molecular Structures through Learnable Atom-Level Watermarking
- Bounded PCTL Model Checking of Large Language Model Outputs
- OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
- VidCLearn: A Continual Learning Approach for Text-to-Video Generation
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- Ensembling Large Language Models for Code Vulnerability Detection: An Empirical Evaluation
- MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Effectively obtaining acoustic, visual and textual data from videos
- Painting the market: generative diffusion models for financial limit order book simulation and forecasting
- TeRA: Rethinking Text-guided Realistic 3D Avatar Generation
- FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework
- Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
- Better Supervised Fine-tuning for VQA: Integer-Only Loss
- AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
- Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Semantic Alignment
- A Meta-Autoethnography of Metadiscourse: Methodological Implications for Interdisciplinary Qualitative Research on Generative Artificial Intelligence Models
- LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- Low-Cost Test-Time Adaptation for Robust Video Editing
- DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
- World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation
- Upsample What Matters: Region-Adaptive Latent Sampling for Accelerated Diffusion Transformers
Discussions
- Sora: Review on Background, Tech, Limits, and Opportunities of Vision Models [hn, 33 points, 2 comments]
- arxiv.org/abs/2402.17177 [bsky, 0 points, 0 comments]
- "This technology opens up possibilities for a more dynamic and interactive form of script development, where ideas can be visualized and assessed in real time, providing a powerful tool for creativity [bsky, 0 points, 1 comments]
Related