AD2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
2025/06/11 by Chenhui Qiang, Wei, Zhaoyang, Qiang, Chenhui +8 · 1 citation
Computer Science · Engineering · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2506.09557
openalex publication_date 2025/06/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Chain-of-Thought (CoT) reasoning has emerged as a powerful approach to enhance the structured, multi-step decision-making capabilities of Multi-Modal Large Models (MLLMs), is particularly crucial for autonomous driving with adverse weather conditions and complex traffic environments. However, existing benchmarks have largely overlooked the need for rigorous evaluation of CoT processes in these specific and challenging scenarios. To address this critical gap, we introduce AD2-Bench, the first Chain-of-Thought benchmark specifically designed for autonomous driving with adverse weather and complex scenes. AD2-Bench is meticulously constructed to fulfill three key criteria: comprehensive data coverage across diverse adverse environments, fine-grained annotations that support multi-step reasoning, and a dedicated evaluation framework tailored for assessing CoT performance. The core contribution of AD2-Bench is its extensive collection of over 5.4k high-quality, manually annotated CoT instances. Each intermediate reasoning step in these annotations is treated as an atomic unit with explicit ground truth, enabling unprecedented fine-grained analysis of MLLMs' inferential processes under text-level, point-level, and region-level visual prompts. Our comprehensive evaluation of state-of-the-art MLLMs on AD2-Bench reveals accuracy below 60%, highlighting the benchmark's difficulty and the need to advance robust, interpretable end-to-end autonomous driving systems. AD2-Bench thus provides a standardized evaluation platform, driving research forward by improving MLLMs' reasoning in autonomous driving, making it an invaluable resource.
Citations
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
- Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
- OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
- Qwen2.5-VL Technical Report
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
- VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
- AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
- Qwen2.5 Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- A-BDD: Leveraging Data Augmentations for Safe Autonomous Driving in Adverse Weather and Lighting
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
- Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- DriveLM: Driving with Graph Visual Question Answering
- LingoQA: Visual Question Answering for Autonomous Driving
- CogAgent: A Visual Language Model for GUI Agents
- NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets using Markup Annotations
- Dolphins: Multimodal Language Model for Driving
- Octopus: Embodied Vision-Language Programmer from Environmental Feedback
- Improved Baselines with Visual Instruction Tuning
- DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
- DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
- Language Prompt for Autonomous Driving
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Drive Like a Human: Rethinking Autonomous Driving with Large Language Models
- NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
- Visual Instruction Tuning
- Open-World Object Manipulation using Pre-trained Vision-Language Models
- VoxFormer: Sparse Voxel Transformer for Camera-based 3D Semantic Scene Completion
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- DRAMA: Joint Risk Localization and Captioning in Driving
- CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving
- Vision-based Large-scale 3D Semantic Mapping for Autonomous Driving Applications
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D
- Explainable Object-induced Action Decision for Autonomous Vehicles
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Argoverse: 3D Tracking and Forecasting with Rich Maps
- nuScenes: A multimodal dataset for autonomous driving
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- Cityscapes dataset for semantic urban scene understanding
Cited by
Related