LaViPlan : Language-Guided Visual Path Planning with RLVR
2025/07/17 by Hoo Oh, Oh, Hayeon
Computer Science · #Advanced Image and Video Retrieval Techniques #Automated planning and scheduling #FOS: Computer and information sciences #Fidelity #Machine Learning (cs.LG) #Motion planning #Multimodal Machine Learning Applications #Qualitative reasoning #Reinforcement learning #Robotic Path Planning Algorithms #Robotics (cs.RO) #Sampling (signal processing) #Spatial intelligence #Verifiable secret sharing
paper · pdf · doi:10.48550/arxiv.2507.12911
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/07/17 · openalex created_date 2025/10/18 · openalex updated_date 2026/08/05
Abstract
Out-of-distribution (OOD) scenarios in autonomous driving pose critical challenges, as planners often fail to generalize beyond their training experience, leading to unsafe or unexpected behavior. Vision-Language Models (VLMs) have shown promise in handling such scenarios by providing high-level scene understanding and user-aligned decisions. However, existing VLMs often exhibit a misalignment between their language-based reasoning and the low-level trajectories required for action-level planning. In this paper, we propose LaViPlan, a framework that leverages Reinforcement Learning with Verifiable Rewards (RLVR) to fine-tune VLMs using planning-oriented metrics. Experimental results show that LaViPlan improves planning performance across both in-domain and out-of-domain datasets. While linguistic fidelity slightly decreases after RLVR-based fine-tuning, qualitative evaluation indicates that the outputs remain coherent. We also conduct ablation studies to analyze the effects of sampling ratio and reasoning guidance, highlighting how these design choices influence performance. These findings demonstrate the potential of RLVR as a post-training paradigm for aligning language-guided reasoning with action-level planning in autonomous driving.
Citations
- RLVR-World: Training World Models with Reinforcement Learning
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
- ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
- Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
- SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
- AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
- Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
- Visual-RFT: Visual Reinforcement Fine-Tuning
- What is the Alignment Objective of GRPO?
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
- Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
- Improving Agent Behaviors with RL Fine-tuning for Autonomous Driving
- CarLLaVA: Vision language models for camera-only closed-loop driving
- ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work Zones
- Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving
- OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
- Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- DriveLM: Driving with Graph Visual Question Answering
- LMDrive: Closed-Loop End-to-End Driving with Large Language Models
- DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Visual Instruction Tuning
- Aligning Text-to-Image Models using Human Feedback
- ADAPT: Action-aware Driving Caption Transformer
- MENLI: Robust Evaluation Metrics from Natural Language Inference
- CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving
- Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration
- BERTScore: Evaluating Text Generation with BERT
- Proximal Policy Optimization Algorithms
Related