World Models Should Prioritize the Unification of Physical and Social Dynamics
2025/10/24 by Zhang, Xiaoyuan, Ma, Chengdong, Huang, Yizhe +5 · 1 citation
#Computers and Society (cs.CY) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.21219
Abstract
World models, which explicitly learn environmental dynamics to lay the foundation for planning, reasoning, and decision-making, are rapidly advancing in predicting both physical dynamics and aspects of social behavior, yet predominantly in separate silos. This division results in a systemic failure to model the crucial interplay between physical environments and social constructs, rendering current models fundamentally incapable of adequately addressing the true complexity of real-world systems where physical and social realities are inextricably intertwined. This position paper argues that the systematic, bidirectional unification of physical and social predictive capabilities is the next crucial frontier for world model development. We contend that comprehensive world models must holistically integrate objective physical laws with the subjective, evolving, and context-dependent nature of social dynamics. Such unification is paramount for AI to robustly navigate complex real-world challenges and achieve more generalizable intelligence. This paper substantiates this imperative by analyzing core impediments to integration, proposing foundational guiding principles (ACE Principles), and outlining a conceptual framework alongside a research roadmap towards truly holistic world models.
Citations
- Social World Model-Augmented Mechanism Design Policy Learning
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
- Embodied AI: From LLMs to World Models
- Empowering Multi-Robot Cooperation via Sequential World Models
- 3D and 4D World Modeling: A Survey
- Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Matrix-3D: Omnidirectional Explorable 3D World Generation
- UserBench: An Interactive Gym Environment for User-Centric Agents
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
- Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
- Back to the Features: DINO as a Foundation for Video World Models
- Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges
- Critique of World Model
- Embodied AI Agents: Modeling the World
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Video World Models with Long-term Spatial Memory
- VRAG: Learning World Models for Interactive Video Generation
- Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
- Learning 3D Persistent Embodied World Models
- SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
- An Evaluation of Cultural Value Alignment in LLM
- LLM Social Simulations Are a Promising Research Method
- World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a Child
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
- Four Principles for Physically Interpretable World Models
- Differentiable Information Enhanced Model-Based Reinforcement Learning
- WorldModelBench: Judging Video Generation Models As World Models
- EgoNormia: Benchmarking Physical Social Norm Understanding
- Intuitive physics understanding emerges from self-supervised pretraining on natural videos
- Improving Transformer World Models for Data-Efficient RL
- A Survey of World Models for Autonomous Driving
- GAWM: Global-Aware World Model for Multi-Agent Reinforcement Learning
- Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition
- Cosmos World Foundation Model Platform for Physical AI
- Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning
- GaussianWorld: Gaussian World Model for Streaming 3D Occupancy\n Prediction
- Navigation World Models
- Open-Sora Plan: Open-Source Large Video Generation Model
- Understanding World or Predicting Future? A Comprehensive Survey of World Models
- AdaSociety: An Adaptive Environment with Social Structures for Multi-Agent Decision-Making
- How Far is Video Generation from World Model: A Physical Law Perspective
- Project Sid: Many-agent simulations toward AI civilization
- MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
- Strong and weak alignment of large language models with human values
- Video Diffusion Alignment via Reward Gradients
- From Efficient Multimodal Models to World Models: A Survey
- Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models
- A Notion of Complexity for Theory of Mind via Discrete World Models
- Pandora: Towards General World Model with Natural Language Actions and Video States
- Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning
- Agent Planning with World Knowledge Model
- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
- Mastering Memory Tasks with World Models
- World Models for Autonomous Driving: An Initial Survey
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
- CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making Agents
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- TD-MPC2: Scalable, Robust World Models for Continuous Control
- DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- Learning to Act from Actionless Videos through Dense Correspondences
- Facing Off World Model Backbones: RNNs, Transformers, and S4
- Training Socially Aligned Language Models on Simulated Social Interactions
- Reasoning with Language Model is Planning with World Model
- Generative Agents: Interactive Simulacra of Human Behavior
- Models as Agents: Optimizing Multi-Step Predictions of Interactive Local Models in Model-Based Multi-Agent Reinforcement Learning
- Mastering Diverse Domains through World Models
- Scalable Diffusion Models with Transformers
- Transformers are Sample-Efficient World Models
- DayDreamer: World Models for Physical Robot Learning
- MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
- Model-based Multi-agent Reinforcement Learning: Recent Progress and Prospects
- Temporal Difference Learning for Model Predictive Control
- Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning
- Mastering Atari with Discrete World Models
- MOPO: Model-based Offline Policy Optimization
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
- Dream to Control: Learning Behaviors by Latent Imagination
- Causality for Machine Learning
- Mastering Atari, Go, chess and shogi by planning with a learned model
- On the Utility of Learning about Humans for Human-AI Coordination
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- When to Trust Your Model: Model-Based Policy Optimization
- Fair Resource Allocation in Federated Learning
- Habitat: A Platform for Embodied AI Research
- Neural MMO: A Massively Multiagent Game Environment for Training and Evaluating Intelligent Agents
- Model-Based Reinforcement Learning for Atari
- Theory of Minds: Understanding Behavior in Groups Through Inverse\n Planning
- M3RL: Mind-aware Multi-agent Management Reinforcement Learning
- World Models
- Machine Theory of Mind
- Compositional Attention Networks for Machine Reasoning
- DeepMind Control Suite
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- CARLA: An Open Urban Driving Simulator
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary\n Visual Reasoning
- Catastrophic Cascade of Failures in Interdependent Networks
- Catastrophic cascade of failures in interdependent networks
- General agents contain world models
- 3D-LLM: Injecting the 3D World into Large Language Models
Cited by
Related