DriveLM: Driving with Graph Visual Question Answering
2023/12/21 by Chonghao Sima, Katrin Renz, Sima, Chonghao +16 · 87 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Graph Neural Networks #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2312.14150
Abstract
We study how vision-language models (VLMs) trained on web-scale data can be integrated into end-to-end driving systems to boost generalization and enable interactivity with human users. While recent approaches adapt VLMs to driving via single-round visual question answering (VQA), human drivers reason about decisions in multiple steps. Starting from the localization of key objects, humans estimate object interactions before taking actions. The key insight is that with our proposed task, Graph VQA, where we model graph-structured reasoning through perception, prediction and planning question-answer pairs, we obtain a suitable proxy task to mimic the human reasoning process. We instantiate datasets (DriveLM-Data) built upon nuScenes and CARLA, and propose a VLM-based baseline approach (DriveLM-Agent) for jointly performing Graph VQA and end-to-end driving. The experiments demonstrate that Graph VQA provides a simple, principled framework for reasoning about a driving scene, and DriveLM-Data provides a challenging benchmark for this task. Our DriveLM-Agent baseline performs end-to-end autonomous driving competitively in comparison to state-of-the-art driving-specific architectures. Notably, its benefits are pronounced when it is evaluated zero-shot on unseen objects or sensor configurations. We hope this work can be the starting point to shed new light on how to apply VLMs for autonomous driving. To facilitate future research, all code, data, and models are available to the public.
Cited by
- World Engine: Towards the Era of Post-Training for Autonomous Driving
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
- ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
- Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
- Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
- FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning
- BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving
- dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- nuScenes Revisited: Progress and Challenges in Autonomous Driving
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- RoboDriveVLM: A Novel Benchmark and Baseline towards Robust Vision-Language Models for Autonomous Driving
- RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models
- Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- FSDAM: Few-Shot Driving Attention Modeling via Vision-Language Coupling
- Argus: Resilience-Oriented Safety Assurance Framework for End-to-End ADSs
- SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World
- AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- SAFe-Copilot: Unified Shared Autonomy Framework
- 3EED: Ground Everything Everywhere in 3D
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
- VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
- SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
- SAVANT: Semantic Analysis with Vision-Augmented Anomaly deTection
- FineVision: Open Data Is All You Need
- SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
- Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
- DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
- Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis
- Align2Act: Instruction-Tuned Models for Human-Aligned Autonomous Driving
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- SanDRA: Safe Large-Language-Model-Based Decision Making for Automated Vehicles Using Reachability Analysis
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- CHAI: Command Hijacking against embodied AI
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- DriveE2E: Closed-Loop Benchmark for End-to-End Autonomous Driving through Real-to-Simulation
- From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving
- BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
- Self-driving cars: Are we there yet?
- MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
- AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving
- The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
- Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning
- Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey
- MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
- OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
- 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model
- CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- DriveQA: Passing the Driving Knowledge Test
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
- Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving
- Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
- SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation
- BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving
Related