WoW: Towards a World omniscient World model Through Embodied Interaction
2025/09/26 by Chi, Xiaowei, Jia, Peidong, Fan, Chun-Kai +33 · 9 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2509.22642
Abstract
Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping physical causality. This observation leads to our central hypothesis: authentic physical intuition of the world model must be grounded in extensive, causally rich interactions with the real world. To test this hypothesis, we present WoW, a 14-billion-parameter generative world model trained on 2 million robot interaction trajectories. Our findings reveal that the model's understanding of physics is a probabilistic distribution of plausible outcomes, leading to stochastic instabilities and physical hallucinations. Furthermore, we demonstrate that this emergent capability can be actively constrained toward physical realism by SOPHIA, where vision-language model agents evaluate the DiT-generated output and guide its refinement by iteratively evolving the language instructions. In addition, a co-trained Inverse Dynamics Model translates these refined plans into executable robotic actions, thus closing the imagination-to-action loop. We establish WoWBench, a new benchmark focused on physical consistency and causal reasoning in video, where WoW achieves state-of-the-art performance in both human and autonomous evaluation, demonstrating strong ability in physical causality, collision dynamics, and object permanence. Our work provides systematic evidence that large-scale, real-world interaction is a cornerstone for developing physical intuition in AI. Models, data, and benchmarks will be open-sourced.
Citations
- ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
- Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Test-time Prompt Refinement for Text-to-Image Models
- AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation
- MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
- Critique of World Model
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
- Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
- Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
- Learning 3D Persistent Embodied World Models
- WorldScore: A Unified Evaluation Benchmark for World Generation
- Segment Any Motion in Videos
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
- Wan: Open and Advanced Large-Scale Video Generative Models
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
- VGGT: Visual Geometry Grounded Transformer
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
- Light-A-Video: Training-free Video Relighting via Progressive Light Fusion
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
- EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
- Do generative video models understand physical principles?
- Cosmos World Foundation Model Platform for Physical AI
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
- MC-LLaVA: Multi-Concept Personalized Vision-Language Model
- MC-LLaVA: Multi-Concept Personalized Vision-Language Model
- RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
- SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
- EVA: An Embodied World Model for Future Video Anticipation
- VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
- VideoAgent: Self-Improving Video Generation
- Lost in Time: A New Temporal Benchmark for VideoLLMs
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Prover-Verifier Games improve legibility of LLM outputs
- LLM Critics Help Catch LLM Bugs
- OpenVLA: An Open-Source Vision-Language-Action Model
- TextGrad: Automatic "Differentiation" via Text
- Unveiling the Tapestry of Consistency in Large Vision-Language Models
- LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
- Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- TempCompass: Do Video LLMs Really Understand Videos?
- Genie: Generative Interactive Environments
- Revisiting Feature Prediction for Learning Visual Representations from Video
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
- EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
- System 2 Attention (is something you might need too)
- Learning to Act from Actionless Videos through Dense Correspondences
- GAIA-1: A Generative World Model for Autonomous Driving
- UniSim: A Neural Closed-Loop Sensor Simulator
- Large Language Models as Commonsense Knowledge for Large-Scale Task Planning
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Automatic Prompt Optimization with "Gradient Descent" and Beam Search
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- DINOv2: Learning Robust Visual Features without Supervision
- Persistent Nature: A Generative Model of Unbounded 3D Worlds
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Diffusion policy: Visuomotor policy learning via action diffusion
- Mastering Diverse Domains through World Models
- Scalable Diffusion Models with Transformers
- Learning Transferable Visual Models From Natural Language Supervision
- Dream to Control: Learning Behaviors by Latent Imagination
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- Learning Latent Dynamics for Planning from Pixels
- World Models
- Understanding deep learning requires rethinking generalization
- A Formal Evaluation of PSNR as Quality Measurement Parameter for Image Segmentation Algorithms
- Euclidean Distance Matrices: Essential theory, algorithms, and applications
- Image quality assessment: from error visibility to structural similarity
Cited by
Related