DriveLM: Driving with Graph Visual Question Answering
2023/12/21 by Chonghao Sima, Katrin Renz, Sima, Chonghao +16 · 167 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Graph Neural Networks #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2312.14150
Abstract
We study how vision-language models (VLMs) trained on web-scale data can be integrated into end-to-end driving systems to boost generalization and enable interactivity with human users. While recent approaches adapt VLMs to driving via single-round visual question answering (VQA), human drivers reason about decisions in multiple steps. Starting from the localization of key objects, humans estimate object interactions before taking actions. The key insight is that with our proposed task, Graph VQA, where we model graph-structured reasoning through perception, prediction and planning question-answer pairs, we obtain a suitable proxy task to mimic the human reasoning process. We instantiate datasets (DriveLM-Data) built upon nuScenes and CARLA, and propose a VLM-based baseline approach (DriveLM-Agent) for jointly performing Graph VQA and end-to-end driving. The experiments demonstrate that Graph VQA provides a simple, principled framework for reasoning about a driving scene, and DriveLM-Data provides a challenging benchmark for this task. Our DriveLM-Agent baseline performs end-to-end autonomous driving competitively in comparison to state-of-the-art driving-specific architectures. Notably, its benefits are pronounced when it is evaluated zero-shot on unseen objects or sensor configurations. We hope this work can be the starting point to shed new light on how to apply VLMs for autonomous driving. To facilitate future research, all code, data, and models are available to the public.
Cited by
- World Engine: Towards the Era of Post-Training for Autonomous Driving
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
- ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
- Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
- Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
- FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning
- BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving
- dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- nuScenes Revisited: Progress and Challenges in Autonomous Driving
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- RoboDriveVLM: A Novel Benchmark and Baseline towards Robust Vision-Language Models for Autonomous Driving
- RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models
- Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- FSDAM: Few-Shot Driving Attention Modeling via Vision-Language Coupling
- Argus: Resilience-Oriented Safety Assurance Framework for End-to-End ADSs
- SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World
- AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- SAFe-Copilot: Unified Shared Autonomy Framework
- 3EED: Ground Everything Everywhere in 3D
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
- RoCA: Robust Cross-Domain End-to-End Autonomous Driving
- VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
- SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
- Can VLMs Unlock Semantic Anomaly Detection? A Framework for Structured Reasoning
- FineVision: Open Data Is All You Need
- SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
- Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
- DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
- Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis
- Align2Act: Instruction-Tuned Models for Human-Aligned Autonomous Driving
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- SanDRA: Safe Large-Language-Model-Based Decision Making for Automated Vehicles Using Reachability Analysis
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- CHAI: Command Hijacking against embodied AI
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- DriveE2E: Closed-Loop Benchmark for End-to-End Autonomous Driving through Real-to-Simulation
- From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving
- BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
- Self-driving cars: Are we there yet?
- MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
- AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving
- The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
- Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning
- Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey
- MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
- OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
- 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model
- CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- DriveQA: Passing the Driving Knowledge Test
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
- ReSim: Reliable World Simulation for Autonomous Driving
- Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving
- Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
- SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation
- BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving
- FedVLM: Scalable Personalized Vision-Language Models through Federated Learning
- Why Braking? Scenario Extraction and Reasoning Utilizing LLM
- AD2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
- LaViPlan : Language-Guided Visual Path Planning with RLVR
- ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
- MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
- Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving
- VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
- Synergistic Prompting for Robust Visual Recognition with Missing Modalities
- LeAD: The LLM Enhanced Planning System Converged with End-to-end Autonomous Driving
- VLAD: A VLM-Augmented Autonomous Driving Framework with Hierarchical Planning and Interpretable Decision Process
- World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
- Automated Vehicles Should be Connected with Natural Language
- Where, What, Why: Towards Explainable Driver Attention Prediction
- DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
- ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
- LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments
- A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects
- Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
- Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review
- Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
- DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
- Demystifying the Visual Quality Paradox in Multimodal Large Language Models
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- NetRoller: Interfacing General and Specialized Models for End-to-End Autonomous Driving
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
- Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
- Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
- Generalized Trajectory Scoring for End-to-end Multimodal Planning
- DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models
- STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
- MLLM-CL: Continual Learning for Multimodal Large Language Models
- Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
- S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
- Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
- GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving
- DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
- RefAV: Towards Planning-Centric Scenario Mining
- Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
- Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
- SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
- DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
- MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- VERDI: VLM-Embedded Reasoning for Autonomous Driving
- TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
- AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
- GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching
- STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision
- Scene-Adaptive Motion Planning with Explicit Mixture of Experts and Interaction-Oriented Optimization
- FIGhost: Fluorescent Ink-based Stealthy and Flexible Backdoor Attacks on Physical Traffic Sign Recognition
- LLM-Assisted Coalition Formation for Cooperative Perception in Autonomous Driving
- Sage Deer: A Super-Aligned Driving Generalist Is Your Copilot
- Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning
- From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
- DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving
- UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
- Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy Scenarios
- Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs
- Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
- χ0: Resource-Aware Robust Manipulation via Taming Distributional Inconsistencies
- Toward Fully Autonomous Driving: AI, Challenges, Opportunities, and Needs
- When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit
- OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
- Planning Safety Trajectories with Dual-Phase, Physics-Informed, and Transportation Knowledge-Driven Large Language Models
- Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
- ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
- Detect Anything 3D in the Wild
- NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
Related