GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
2025/03/26 by Russell, Lloyd, Hu, Anthony, Bertoni, Lorenzo +4 · 44 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2503.20523
Abstract
Generative models offer a scalable and flexible paradigm for simulating complex environments, yet current approaches fall short in addressing the domain-specific requirements of autonomous driving - such as multi-agent interactions, fine-grained control, and multi-camera consistency. We introduce GAIA-2, Generative AI for Autonomy, a latent diffusion world model that unifies these capabilities within a single generative framework. GAIA-2 supports controllable video generation conditioned on a rich set of structured inputs: ego-vehicle dynamics, agent configurations, environmental factors, and road semantics. It generates high-resolution, spatiotemporally consistent multi-camera videos across geographically diverse driving environments (UK, US, Germany). The model integrates both structured conditioning and external latent embeddings (e.g., from a proprietary driving model) to facilitate flexible and semantically grounded scene synthesis. Through this integration, GAIA-2 enables scalable simulation of both common and rare driving scenarios, advancing the use of generative world models as a core tool in the development of autonomous systems. Videos are available at https://wayve.ai/thinking/gaia-2.
Cited by
- AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models
- Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space
- AstraNav-World: World Model for Foresight Control and Consistency
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
- Dexterous World Models
- RadarGen: Automotive Radar Point Cloud Generation from Cameras
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
- Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- IC-World: In-Context Generation for Shared World Modeling
- What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
- GigaWorld-0: World Models as Data Engine to Empower Embodied AI
- DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- Driving scenario generation and evaluation using a structured layer representation and foundational models
- CG-World: A Large-Scale World-State Dataset and Protocol for World Models
- Generative View Stitching
- AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- Learning to Generate Rigid Body Interactions with Video Diffusion Models
- Training Agents Inside of Scalable World Models
- From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- MAD: Motion Appearance Decoupling for efficient Driving World Models
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- Back to the Features: DINO as a Foundation for Video World Models
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- A Concept for Efficient Scalability of Automated Driving Allowing for Technical, Legal, Cultural, and Ethical Differences
- HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study
- Controllable Video Generation: A Survey
- AirScape: An Aerial Generative World Model with Motion Controllability
Related