WorldMark: A Unified Benchmark Suite for Interactive Video World Models
2026/04/30 by Xiaojie Xu, Zhengyuan Lin, Kang He +5
Computer Science · #cs.CV
paper · pdf · doi:10.48550/arxiv.2604.21686
arxiv created 2026/08/05 · arxiv updated 2026/08/06
Abstract
Unlike text- or image-driven video generation, an interactive world model is driven by actions: the user acts, and the world responds. Two obstacles stand in the way of fair and comprehensive evaluation. First, models take actions in incompatible formats---captions, camera trajectories, action functions---so no shared protocol has been established. Second, while existing benchmarks have advanced world memory and visual quality, action following is reduced to trajectory or direction error, which collapses a whole path into one number: not how quickly the world reacts to a command switch, nor how cleanly it moves along the commanded axis. WorldMark removes both obstacles. Per-model adapters translate a shared WASD-style vocabulary into each model's native control format, so ten heterogeneous models receive semantically identical instructions across 500 standardized cases spanning styles, viewpoints, and difficulty tiers; a new model costs one adapter. On this common ground we characterize action dynamics through a control-systems lens---direction accuracy, direction purity, response latency, and motion stability, each resolved per axis---alongside suites for world memory and visual quality. Together they expose differences existing protocols cannot see: the fastest responders are often the least stable, a trade-off no single action metric captures; per-axis resolution reveals models that follow translation almost perfectly while barely responding to rotation; the model with the best perceptual and aesthetic quality ranks last in translational direction accuracy and latency; and stylized scenes cost every model global consistency while leaving action dynamics largely intact. We will release all data, evaluation code, and model outputs.
Citations
- DreamX-World 1.0: A General-Purpose Interactive World Model
- SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
- Lyra 2.0: Explorable Generative 3D Worlds
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- Yume-1.5: A Text-Controlled Interactive World Generation Model
- SVBench: Evaluation of Video Generation Models on Social Reasoning
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
- Depth Anything 3: Recovering the Visual Space from Any Views
- Matrix-game 2.0: An open-source real-time and streaming interactive world model
- ViPE: Video Pose Engine for 3D Geometric Perception
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
- Yume: An Interactive World Generation Model
- π3: Permutation-Equivariant Visual Geometry Learning
- WorldVLA: Towards Autoregressive Action World Model
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- WorldScore: A Unified Evaluation Benchmark for World Generation
- Wan: Open and Advanced Large-Scale Video Generative Models
- WorldModelBench: Judging Video Generation Models As World Models
- GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
- Diffusion Models Are Real-Time Game Engines
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- SEA-RAFT: Simple, Efficient, Accurate RAFT for Optical Flow
- Diffusion for World Modeling: Visual Details Matter in Atari
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- Genie: Generative Interactive Environments
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
- MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Learning Interactive Real-World Simulators
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- DINOv2: Learning Robust Visual Features without Supervision
- Learning-based Inverse Rendering of Complex Indoor Scenes with Differentiable Monte Carlo Raytracing
- Video Diffusion Models
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- MUSIQ: Multi-scale Image Quality Transformer
- Aligning Latent and Image Spaces to Connect the Unconnectable
- EDEN: Multimodal Synthetic Dataset of Enclosed GarDEN Scenes
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
- TransNet V2: An effective deep network architecture for fast shot transition detection
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- World Models
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Matrix-Game: Interactive World Foundation Model