CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
2025/11/25 by Zhang, Dapeng, Shen, Fei, Zhao, Rui +5
#FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2511.19914
Abstract
Autonomous driving represents a prominent application of artificial intelligence. Recent approaches have shifted from focusing solely on common scenarios to addressing complex, long-tail situations such as subtle human behaviors, traffic accidents, and non-compliant driving patterns. Given the demonstrated capabilities of large language models (LLMs) in understanding visual and natural language inputs and following instructions, recent methods have integrated LLMs into autonomous driving systems to enhance reasoning, interpretability, and performance across diverse scenarios. However, existing methods typically rely either on real-world data, which is suitable for industrial deployment, or on simulation data tailored to rare or hard case scenarios. Few approaches effectively integrate the complementary advantages of both data sources. To address this limitation, we propose a novel VLM-guided, end-to-end adversarial transfer framework for autonomous driving that transfers long-tail handling capabilities from simulation to real-world deployment, named CoC-VLA. The framework comprises a teacher VLM model, a student VLM model, and a discriminator. Both the teacher and student VLM models utilize a shared base architecture, termed the Chain-of-Causality Visual-Language Model (CoC VLM), which integrates temporal information via an end-to-end text adapter. This architecture supports chain-of-thought reasoning to infer complex driving logic. The teacher and student VLM models are pre-trained separately on simulated and real-world datasets. The discriminator is trained adversarially to facilitate the transfer of long-tail handling capabilities from simulated to real-world environments by the student VLM model, using a novel backpropagation strategy.
Citations
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- AutoDrive-R2: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
- EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning
- DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo
- AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
- A Dual-Agent Adversarial Framework for Robust Generalization in Deep Reinforcement Learning
- Patch-GAN Transfer Learning with Reconstructive Models for Cloud Removal
- DTSGAN: Learning Dynamic Textures via Spatiotemporal Generative Adversarial Network
- MapExpert: Online HD Map Construction with Simple and Efficient Sparse Map Element Expert
- DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View Stereo
- LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
- ContextVLM: Zero-Shot and Few-Shot Context Understanding for Autonomous Driving using Vision Language Models
- MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View Stereo
- SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
- TokenUnify: Scaling Up Autoregressive Pretraining for Neuron Segmentation
- NeuroNCAP: Photorealistic Closed-loop Safety Testing for Autonomous Driving
- BIMCV-R: A Landmark Dataset for 3D CT Text-Image Retrieval
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- DriveLM: Driving with Graph Visual Question Answering
- LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding
- DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
- LMDrive: Closed-Loop End-to-End Driving with Large Language Models
- Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
- NeuRAD: Neural Rendering for Autonomous Driving
- DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
- FusionAD: Multi-modality Fusion for Prediction and Planning Tasks of Autonomous Driving
- DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving
- ReasonNet: End-to-End Driving with Temporal and Global Reasoning
- Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes
- Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- VAD: Vectorized Scene Representation for Efficient Autonomous Driving
- LLaMA: Open and Efficient Foundation Language Models
- GPTScore: Evaluate as You Desire
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Planning-oriented Autonomous Driving
- Perceive, Interact, Predict: Learning Dynamic and Static Clues for End-to-End Motion Prediction
- MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
- ViP3D: End-to-end Visual Trajectory Prediction via 3D Agent Queries
- ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning
- HOPE: Hierarchical Spatial-temporal Network for Occupancy Flow Prediction
- VectorMapNet: End-to-end Vectorized HD Map Learning
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline
- PETR: Position Embedding Transformation for Multi-View 3D Object\n Detection
- GRI: General Reinforced Imitation and its Application to Vision-Based Autonomous Driving
- DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries
- Generative Adversarial Networks
- End-to-End Urban Driving by Imitating a Reinforcement Learning Coach
- HDMapNet: An Online HD Map Construction and Evaluation Framework
- LoRA: Low-Rank Adaptation of Large Language Models
- Scene Transformer: A unified architecture for predicting multiple agent trajectories
- Learning to drive from a world on rails
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular Cameras
- Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
- MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain Adaptation
- Multimodal Motion Prediction with Stacked Transformers
- Learning Transferable Visual Models From Natural Language Supervision
- MP3: A Unified Model to Map, Perceive, Predict and Plan
- Perceive, Predict, and Plan: Safe Motion Planning Through Interpretable\n Semantic Representations
- Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by\n Implicitly Unprojecting to 3D
- PnPNet: End-to-End Perception and Prediction with Tracking in the Loop
- Language Models are Few-Shot Learners
- VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation
- CoverNet: Multimodal Behavior Prediction using Trajectory Sets
- MultiPath: Multiple Probabilistic Anchor Trajectory Hypotheses for Behavior Prediction
- CenterNet: Keypoint Triplets for Object Detection
- nuScenes: A multimodal dataset for autonomous driving
- Using Pre-Training Can Improve Model Robustness and Uncertainty
- PointPillars: Fast Encoders for Object Detection from Point Clouds
- IntentNet: Learning to Predict Intention from Raw Sensor Data
- LaneNet: Real-Time Lane Detection Networks for Autonomous Driving
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- Deep Adversarial Attention Alignment for Unsupervised Domain Adaptation: the Benefit of Target Expectation Maximization
- CARLA: An Open Urban Driving Simulator
- Image-to-Image Translation with Conditional Adversarial Networks
- Domain Separation Networks
- SPICE: Semantic Propositional Image Caption Evaluation
- CIDEr: Consensus-based Image Description Evaluation
- Unsupervised Domain Adaptation by Backpropagation
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
Related