WorldVLA: Towards Autoregressive Action World Model
2025/06/26 by Jun Cen, Cen, Jun, Chaohui Yu +21 · 4 voices · 102 citations
Computer Science · #Action (physics) #Action recognition #Autoregressive model #Data Visualization and Analytics #Generalization #Image (mathematics)
paper · pdf · doi:10.48550/arxiv.2506.21539
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/06/26 · openalex created_date 2025/10/15 · openalex updated_date 2026/08/05
Abstract
We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA intergrates Vision-Language-Action (VLA) model and world model in one single framework. The world model predicts future images by leveraging both action and image understanding, with the purpose of learning the underlying physics of the environment to improve action generation. Meanwhile, the action model generates the subsequent actions based on image observations, aiding in visual understanding and in turn helps visual generation of the world model. We demonstrate that WorldVLA outperforms standalone action and world models, highlighting the mutual enhancement between the world model and the action model. In addition, we find that the performance of the action model deteriorates when generating sequences of actions in an autoregressive manner. This phenomenon can be attributed to the model's limited generalization capability for action prediction, leading to the propagation of errors from earlier actions to subsequent ones. To address this issue, we propose an attention mask strategy that selectively masks prior actions during the generation of the current action, which shows significant performance improvement in the action chunk generation task.
Citations
Cited by
- KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
- Do World Action Models Generalize Better than VLAs? A Robustness Study
- RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
- RhinoVLA Technical Report
- GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
- Native Video-Action Pretraining for Generalizable Robot Control
- FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
- Act2Goal: From World Model To General Goal-conditioned Policy
- WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
- N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
- DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
- CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
- A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models
- AstraNav-World: World Model for Foresight Control and Consistency
- Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
- Active Intelligence in Video Avatars via Closed-loop World Modeling
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Efficient-VLN: A Training-Efficient Vision-Language Navigation Model
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
- Astra: General Interactive World Model with Autoregressive Denoising
- VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- Mixture of Horizons in Action Chunking
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- Towards High-Consistency Embodied World Model with Multi-View Trajectory Videos
- DAP: A Discrete-token Autoregressive Planner for Autonomous Driving
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
- A Step Toward World Models: A Survey on Robotic Manipulation
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
- PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
- Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?
- CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
- A Comprehensive Survey on World Models for Embodied AI
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- QuantWAMs: Calibrating at the Right Granularity for World Action Models
- Self-evolved Imitation Learning in Simulated World
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- MAPF-World: Action World Model for Multi-Agent Path Finding
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- MolmoAct: Action Reasoning Models that can Reason in Space
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
- Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends
- CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
- Hermite Curves as Trajectory Priors for Vision-Language-Action Models
- EndoWAM: A Grounded World-Action Model for Generalizable Endoscopic Navigation
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
- DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- Enhancing Policy Learning with World-Action Model
- Geometric Action Model for Robot Policy Learning
- VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
- Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
- VLANeXt: Recipes for Building Strong VLA Models
- VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
- Learning Action Priors for Cross-embodiment Robot Manipulation
- World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
- StarVLA-α: Reducing Complexity in Vision-Language-Action Systems
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
- CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
- Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models
- DreamWAM: Beyond RGB Future Prediction for World Action Models
- WorldMark: A Unified Benchmark Suite for Interactive Video World Models
- JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment
- Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models
- DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
Discussions
Related