Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
2026/04/20 by Jinghui Lu, Jiayi Guan, Zhijian Huang +47 · 2 voices · 3 citations
Computer Science · #cs.CV #cs.CL #cs.RO
paper · pdf
arxiv published 2026/04/20 · arxiv updated 2026/05/08
Abstract
Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is prohibitive for real-time deployment. Latent CoT methods attempt to close this gap by compressing reasoning into continuous hidden states, but consistently fall short of their explicit counterparts. We suggest that this is due to purely linguistic latent representations compressing a symbolic abstraction of the world, rather than the causal dynamics that actually govern driving. Thus, we present OneVL (One-step latent reasoning and planning with Vision-Language explanations), a unified VLA and World Model framework that routes reasoning through compact latent tokens supervised by dual auxiliary decoders. Alongside a language decoder that reconstructs text CoT, we introduce a visual world model decoder that predicts future-frame tokens, forcing the latent space to internalize the causal dynamics of road geometry, agent motion, and environmental change. A three-stage training pipeline progressively aligns these latents with trajectory, language, and visual objectives, ensuring stable joint optimization. In inference, the auxiliary decoders are discarded, and all latent tokens are prefilled in a single parallel pass, matching the speed of answer-only prediction. Across four benchmarks, OneVL becomes the first latent CoT method to surpass explicit CoT, delivering superior accuracy at answer-only latency. These results show that with world model supervision, latent CoT produces more generalizable representations than verbose token-by-token reasoning. Code has been open-sourced to the community. Project Page: https://xiaomi-embodied-intelligence.github.io/OneVL
Citations
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
- The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving
- DVGT: Driving Visual Geometry Transformer
- MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World
- U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
- SimScale: Learning to Drive via Real-World Simulation at Scale
- Qwen3-VL Technical Report
- AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
- GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
- MiMo-Embodied: X-Embodied Foundation Model Technical Report
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Emu3.5: Native Multimodal Models are World Learners
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
- SIM-CoT: Supervised Implicit Chain-of-Thought
- AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving
- 3D and 4D World Modeling: A Survey
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
- NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
- DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
- ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
- Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
- Advancing Sequential Numerical Prediction in Autoregressive Models
- Seed1.5-VL Technical Report
- End-to-End Driving with Online Trajectory Evaluation via BEV World Model
- Vision as LoRA
- ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
- DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
- Qwen2.5-VL Technical Report
- SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs
- Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
- A Survey of World Models for Autonomous Driving
- LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
- Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
- RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
- Reasoning Beyond Words ? Exploring framework for hidden state reasoning
- Scalable Image Tokenization with Index Backpropagation Quantization
- GaussianPretrain: A Simple Unified 3D Gaussian Representation for Visual Pre-training in Autonomous Driving
- GPT-4o System Card
- DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes
- Making Large Language Models Better Planners with Reasoning-Decision Alignment
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
- NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
- Enhancing End-to-End Autonomous Driving with Latent World Model
- ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
- The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition
- Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
- OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
- Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
- Towards learning-based planning:The nuPlan benchmark for real-world autonomous driving
- World Models for Autonomous Driving: An Initial Survey
- PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity Recognition
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
- DriveLM: Driving with Graph Visual Question Answering
- LingoQA: Visual Question Answering for Autonomous Driving
- Improved Baselines with Visual Instruction Tuning
- Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving
- DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
- GAIA-1: A Generative World Model for Autonomous Driving
- Language Modeling Is Compression
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration
- Let's Verify Step by Step
- Visual Instruction Tuning
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting
- PUnifiedNER: A Prompting-based Unified NER System for Diverse Datasets
- What Makes Pre-trained Language Models Better Zero-shot Learners?
- DRAMA: Joint Risk Localization and Captioning in Driving
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- A Rationale-Centric Framework for Human-in-the-loop Machine Learning
- Evaluating Large Language Models Trained on Code
- One Million Scenes for Autonomous Driving: ONCE Dataset
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset
- Dream to Control: Learning Behaviors by Latent Imagination
- nuScenes: A multimodal dataset for autonomous driving
- IDD: A Dataset for Exploring Problems of Autonomous Navigation in Unconstrained Environments
- World Models
- Attention Is All You Need
- Deep Learning and the Information Bottleneck Principle
- Vision meets robotics: The KITTI dataset
- Universal Intelligence: A Definition of Machine Intelligence
- Universal Intelligence: A Definition of Machine Intelligence
- The information bottleneck method
- OpenAI o1 System Card
Discussions
- [20/30] 182 Upvotes, 4 Comments, 2 Posts, arXiv:2604.18486 🆕OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li, Guang Li, [bsky, 0 points, 2 comments]
- 今日読んだ論文 OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation arxiv.org/abs/2604.18486 Reasoningをテキストで明にやるのではなく、潜在空間で固定長トークンをprefill部分に埋め込んで、そこから学習時だけDecoderでやっていくと。Encoder-De [bsky, 0 points, 0 comments]
Related