OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
2025/09/15 by Yang Zhou, Zhou, Yang, Y. Q. Wang +35 · 8 citations
Engineering · Environmental Science · #3D Modeling in Geospatial Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Remote Sensing and LiDAR Applications
paper · pdf · doi:10.48550/arxiv.2509.12201
openalex publication_date 2025/09/15 · openalex created_date 2025/10/12 · openalex updated_date 2026/07/28
Abstract
The field of 4D world modeling - aiming to jointly capture spatial geometry and temporal dynamics - has witnessed remarkable progress in recent years, driven by advances in large-scale generative models and multimodal learning. However, the development of truly general 4D world models remains fundamentally constrained by the availability of high-quality data. Existing datasets and benchmarks often lack the dynamic complexity, multi-domain diversity, and spatial-temporal annotations required to support key tasks such as 4D geometric reconstruction, future prediction, and camera-control video generation. To address this gap, we introduce OmniWorld, a large-scale, multi-domain, multi-modal dataset specifically designed for 4D world modeling. OmniWorld consists of a newly collected OmniWorld-Game dataset and several curated public datasets spanning diverse domains. Compared with existing synthetic datasets, OmniWorld-Game provides richer modality coverage, larger scale, and more realistic dynamic interactions. Based on this dataset, we establish a challenging benchmark that exposes the limitations of current state-of-the-art (SOTA) approaches in modeling complex 4D environments. Moreover, fine-tuning existing SOTA methods on OmniWorld leads to significant performance gains across 4D reconstruction and video generation tasks, strongly validating OmniWorld as a powerful resource for training and evaluation. We envision OmniWorld as a catalyst for accelerating the development of general-purpose 4D world models, ultimately advancing machines' holistic understanding of the physical world.
Citations
- Matrix-game 2.0: An open-source real-time and streaming interactive world model
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
- Sekai: A Video Dataset towards World Exploration
- Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
- DeepVerse: 4D Autoregressive Video Generation as a World Model
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- Dynamic Camera Poses and Where to Find Them
- Aether: Geometric-Aware Unified World Modeling
- RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation
- DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Framework
- VGGT: Visual Geometry Grounded Transformer
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
- FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Continuous 3D Perception Model with Persistent State
- FoundationStereo: Zero-Shot Stereo Matching
- GameFactory: Creating New Games with Generative Interactive Videos
- Cosmos World Foundation Model Platform for Physical AI
- Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization
- MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds
- MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
- The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control
- AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
- MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
- CamI2V: Camera-Controlled Image-to-Video Diffusion Model
- Depth Any Video with Scalable Synthetic Data
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- SAM 2: Segment Anything in Images and Videos
- MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
- Grounding Image Matching in 3D with MASt3R
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
- RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos
- DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision
- DUSt3R: Geometric 3D Vision Made Easy
- MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
- DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World
- ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking
- RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot
- DynamicStereo: Consistent Dynamic Depth from Stereo Videos
- DataComp: In search of the next generation of multimodal datasets
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Adding Conditional Control to Text-to-Image Diffusion Models
- Adding Conditional Control to Text-to-Image Diffusion Models
- Mastering Diverse Domains through World Models
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Particle Video Revisited: Tracking Through Occlusions Using Point Trajectories
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- EDEN: Multimodal Synthetic Dataset of Enclosed GarDEN Scenes
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- BlendedMVS: A Large-scale Dataset for Generalized Multi-view Stereo Networks
- Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera
- ReFusion: 3D Reconstruction in Dynamic Environments for RGB-D Cameras\n Exploiting Residuals
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- MegaDepth: Learning Single-View Depth Prediction from Internet Photos
- World Models
- SuperPoint: Self-Supervised Interest Point Detection and Description
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- Playing for Data: Ground Truth from Computer Games
- Vision meets robotics: The KITTI dataset
- Distinctive Image Features from Scale-Invariant Keypoints
Cited by
Related