Quo Vadis, World Modeling?
2026/08/03 by Yu Yang, Xuemeng Yang, Licheng Wen +17
Computer Science · #cs.CV #cs.AI #cs.RO
paper · pdf
Technical Blog at https://worldbench.github.io/awesome-agentic-world-model GitHub Repo at https://github.com/worldbench/awesome-agentic-world-model
arxiv created 2026/08/03 · arxiv updated 2026/08/05
Abstract
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference-Time Guidance, where proxy outputs enrich in-context information for superior decisions; L.2 Training-Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent-Proxy Co-Evolution, where real-environment evidence continuously updates both the proxy and the agent for co-evolution. Ultimately, this work recasts world modeling into an agent-centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.
Citations
- MemHarness: Memory Is Reconstructed, Not Replayed
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
- The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
- Computer-Using World Model
- WebWorld: A Large-Scale World Model for Web Agent Training
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
- Generative Visual Code Mobile World Models
- Web World Models
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- GTM: Simulating the World of Tools for AI Agents
- U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
- AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
- Computer-Use Agents as Judges for Generative User Interface
- PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Scaling Agent Learning via Experience Synthesis
- VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
- Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models
- A Comprehensive Survey on World Models for Embodied AI
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- Agent Learning via Early Experience
- CWM: An Open-Weights LLM for Research on Code Generation with World Models
- 3D and 4D World Modeling: A Survey
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
- NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
- WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
- ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
- MindCube: Spatial Mental Modeling from Limited Views
- GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- World Modelling Improves Language Model Agents
- Occupancy World Model for Robots
- TesserAct: Learning 4D Embodied World Models
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
- ViMo: A Generative Visual GUI World Model for App Agents
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- Aether: Geometric-Aware Unified World Modeling
- VGGT: Visual Geometry Grounded Transformer
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- WorldModelBench: Judging Video Generation Models As World Models
- The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Continuous 3D Perception Model with Persistent State
- A Survey of World Models for Autonomous Driving
- GameFactory: Creating New Games with Generative Interactive Videos
- Cosmos World Foundation Model Platform for Physical AI
- MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
- Navigation World Models
- Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
- A Survey on LLM-as-a-Judge
- Understanding World or Predicting Future? A Comprehensive Survey of World Models
- Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
- DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
- How Far is Video Generation from World Model: A Physical Law Perspective
- GameGen-X: Interactive Open-world Game Video Generation
- World Models: The Safety Perspective
- DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes
- WorldSimBench: Towards Video Generation Models as World Simulators
- Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
- DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
- 3D Reconstruction with Spatial Memory
- Diffusion Models Are Real-Time Game Engines
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
- LLM Critics Help Catch LLM Bugs
- Grounding Image Matching in 3D with MASt3R
- WonderWorld: Interactive 3D Scene Generation from a Single Image
- Evaluating the World Model Implicit in a Generative Model
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- iVideoGPT: Interactive VideoGPTs are Scalable World Models
- Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Agent Planning with World Knowledge Model
- Diffusion for World Modeling: Visual Details Matter in Atari
- DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving
- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
- ORPO: Monolithic Preference Optimization without Reference Model
- Continual Learning and Catastrophic Forgetting
- Genie: Generative Interactive Environments
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
- World Model on Million-Length Video And Language With Blockwise RingAttention
- CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- KTO: Model Alignment as Prospect Theoretic Optimization
- A Survey on 3D Gaussian Splatting
- DUSt3R: Geometric 3D Vision Made Easy
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
- Mip-Splatting: Alias-free 3D Gaussian Splatting
- OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
- Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion
- TD-MPC2: Scalable, Robust World Models for Continuous Control
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- STORM: Efficient Stochastic Transformer based World Models for Reinforcement Learning
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- MemGPT: Towards LLMs as Operating Systems
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Learning Interactive Real-World Simulators
- GAIA-1: A Generative World Model for Autonomous Driving
- DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
- Cognitive Architectures for Language Agents
- CityDreamer: Compositional Generative Model of Unbounded 3D Cities
- AgentBench: Evaluating LLMs as Agents
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- Autonomous Tester Agent Benchmark
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Let's Verify Step by Step
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Reasoning with Language Model is Planning with World Model
- Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- RRHF: Rank Responses to Align Language Models with Human Feedback without tears
- Generative Agents: Interactive Simulacra of Human Behavior
- Self-Refine: Iterative Refinement with Self-Feedback
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Toolformer: Language Models Can Teach Themselves to Use Tools
- SceneScape: Text-Driven Consistent Scene Generation
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Affective Coherence Monitoring for Transformer-Based Language Models
- RT-1: Robotics Transformer for Real-World Control at Scale
- Solving math word problems with process- and outcome-based feedback
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Code as Policies: Language Model Programs for Embodied Control
- Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
- Transformers are Sample-Efficient World Models
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- DayDreamer: World Models for Physical Robot Learning
- Diffusion Models for Video Prediction and Infilling
- Large Language Models are Zero-Shot Reasoners
- MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Video Diffusion Models
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- TensoRF: Tensorial Radiance Fields
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Instant neural graphics primitives with a multiresolution hash encoding
- Plenoxels: Radiance Fields without Neural Networks
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
- Training Verifiers to Solve Math Word Problems
- CLIPort: What and Where Pathways for Robotic Manipulation
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- FitVid: Overfitting in Pixel-Level Video Prediction
- Pathdreamer: A World Model for Indoor Navigation
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields
- Transporter Networks: Rearranging the Visual World for Robotic Manipulation
- Learning to summarize from human feedback
- Dream to Control: Learning Behaviors by Latent Imagination
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
- Fine-Tuning Language Models from Human Preferences
- When to Trust Your Model: Model-Based Policy Optimization
- Habitat: A Platform for Embodied AI Research
- Model-Based Reinforcement Learning for Atari
- Learning Latent Dynamics for Planning from Pixels
- Recurrent World Models Facilitate Policy Evolution
- Proximal Policy Optimization Algorithms
- Deep reinforcement learning from human preferences
- Matrix-Game: Interactive World Foundation Model