Toward Fully Autonomous Driving: AI, Challenges, Opportunities, and Needs
2026/01/30 by Lars Ullrich, Michael Buchholz, Klaus Dietmayer +1 · 1 voice
Computer Science · #cs.RO #cs.ET
paper · pdf · doi:10.1109/access.2026.3659192
Abstract
Automated driving (AD) is promising, but the transition to fully autonomous driving is, among other things, subject to the real, ever-changing open world and the resulting challenges. However, research in the field of AD demonstrates the ability of artificial intelligence (AI) to outperform classical approaches, handle higher complexities, and reach a new level of autonomy. At the same time, the use of AI raises further questions of safety and transferability. To identify the challenges and opportunities arising from AI concerning autonomous driving functionalities, we have analyzed the current state of AD, outlined limitations, and identified foreseeable technological possibilities. Thereby, various further challenges are examined in the context of prospective developments. In this way, this article reconsiders fully autonomous driving with respect to advancements in the field of AI and carves out the respective needs and resulting research questions.
Citations
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- A New Perspective On AI Safety Through Control Theory Methodologies
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
- V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models
- Let Occ Flow: Self-Supervised 3D Occupancy Flow Prediction
- OpenVLA: An Open-Source Vision-Language-Action Model
- Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
- Fast Collision Probability Estimation for Automated Driving using Multi-circular Shape Approximations
- ViewFormer: Exploring Spatiotemporal Modeling for Multi-View 3D Occupancy Perception via View-Guided Transformers
- Incorporating Explanations into Human-Machine Interfaces for Trust and Situation Awareness in Autonomous Vehicles
- QuAD: Query-based Interpretable Neural Motion Planning for Autonomous Driving
- Safety Implications of Explainable Artificial Intelligence in End-to-End Autonomous Driving
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- GenAD: Generative End-to-End Autonomous Driving
- BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving
- DriveLM: Driving with Graph Visual Question Answering
- Comparison of Waymo Rider-Only Crash Data to Human Benchmarks at 7.1 Million Miles
- DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
- MGTR: Multi-Granular Transformer for Motion Prediction with LiDAR
- LLM4Drive: A Survey of Large Language Models for Autonomous Driving
- Interactive Joint Planning for Autonomous Vehicles
- Vision Language Models in Autonomous Driving: A Survey and Outlook
- DTPP: Differentiable Joint Conditional Prediction and Cost Evaluation for Tree Policy Planning in Autonomous Driving
- LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
- Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving
- DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
- GPT-Driver: Learning to Drive with GPT
- MotionLM: Multi-Agent Motion Forecasting as Language Modeling
- The Integration of Prediction and Planning in Deep Learning Automated Driving Systems: A Review
- FusionAD: Multi-modality Fusion for Prediction and Planning Tasks of Autonomous Driving
- DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
- End-to-end Autonomous Driving: Challenges and Frontiers
- Scene as Occupancy
- Building a Credible Case for Safety: Waymo's Approach for the Determination of Absence of Unreasonable Risk
- Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving
- Connecting the Dots in Trustworthy Artificial Intelligence: From AI Principles, Ethics, and Key Requirements to Responsible AI Systems and Regulation
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy Prediction
- Micrograph segmentations for DDEVD
- Segment Anything
- ChatGPT for good? On opportunities and challenges of large language models for education
- VAD: Vectorized Scene Representation for Efficient Autonomous Driving
- GPT-4 Technical Report
- GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous Driving
- LLaMA: Open and Efficient Foundation Language Models
- Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction
- Tree-structured Policy Planning with Learned Behavior Models
- Planning-oriented Autonomous Driving
- DiffStack: A Differentiable and Modular Control Stack for Autonomous Vehicles
- Safe Real-World Autonomous Driving by Learning to Predict and Plan with a Mixture of Experts
- PlanT: Explainable Planning Transformers via Object-Level Representations
- Generalizing in the Real World with Representation Learning
- Model-Based Imitation Learning for Urban Driving
- Visual Language Maps for Robot Navigation
- Differentiable Raycasting for Self-supervised Occupancy Forecasting
- Motion Transformer with Global Intention Localization and Local Movement Refinement
- ViP3D: End-to-end Visual Trajectory Prediction via 3D Agent Queries
- Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer
- Differentiable Integrated Motion Prediction and Planning with Learnable Cost Function for Autonomous Driving
- ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning
- YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
- YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline
- TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving
- BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation
- HDGT: Heterogeneous Driving Graph Transformer for Multi-Agent Trajectory Prediction via Scene Encoding
- PaLM: Scaling Language Modeling with Pathways
- BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers
- Learning from All Vehicles
- M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction
- Indy Autonomous Challenge -- Autonomous Race Cars at the Handling Limits
- Explainable Artificial Intelligence for Autonomous Driving: A Comprehensive Overview and Field Guide for Future Research Directions
- Masked-attention Mask Transformer for Universal Image Segmentation
- Swin Transformer V2: Scaling Up Capacity and Resolution
- MTP: Multi-Hypothesis Tracking and Prediction for Reduced Error Propagation
- Injecting Planning-Awareness into Prediction and Detection Evaluation
- Mismatched No More: Joint Model-Policy Optimization for Model-Based RL
- SafetyNet: Safe planning for real-world self-driving vehicles using machine-learned policies
- Urban Driver: Learning to Drive from Real-world Demonstrations Using Policy Gradients
- NEAT: Neural Attention Fields for End-to-End Autonomous Driving
- Exploring Simple 3D Multi-Object Tracking for Autonomous Driving
- End-to-End Urban Driving by Imitating a Reinforcement Learning Coach
- Reimagining an autonomous vehicle
- Shifts: A Dataset of Real Distributional Shift Across Multiple Large-Scale Tasks
- NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles
- MOTR: End-to-End Multiple-Object Tracking with Transformer
- Learning to drive from a world on rails
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular Cameras
- Contingencies from Observations: Tractable Contingency Planning with Learned Behavior Models
- Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
- Adaptive Methods for Real-World Domain Generalization
- MP3: A Unified Model to Map, Perceive, Predict and Plan
- Probabilistic 3D Multi-Modal, Multi-Object Tracking for Autonomous Driving
- GPT-3: Its Nature, Scope, Limits, and Consequences
- Computing Systems for Autonomous Driving: State-of-the-Art and Challenges
- Perceive, Predict, and Plan: Safe Motion Planning Through Interpretable Semantic Representations
- Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D
- DSDNet: Deep Structured self-Driving Network
- Implicit Latent Variable Model for Scene-Consistent Motion Forecasting
- DeepCLR: Correspondence-Less Architecture for Deep End-to-End Point Cloud Registration
- 1st Place Solution for Waymo Open Dataset Challenge -- 3D Detection and Domain Adaptation
- Can Autonomous Vehicles Identify, Recover From, and Adapt to Distribution Shifts?
- One Thousand and One Hours: Self-driving Motion Prediction Dataset
- The Importance of Prior Knowledge in Precise Multimodal Prediction
- PnPNet: End-to-End Perception and Prediction with Tracking in the Loop
- End-to-End Object Detection with Transformers
- Reducing Uncertainty by Fusing Dynamic Occupancy Grid Maps in a Cloud-based Collective Environment Model
- Learning to Evaluate Perception Models Using Planner-Centric Metrics
- Meta-Learning in Neural Networks: A Survey
- Tracking Objects as Points
- Tracking Objects as Points
- PiP: Planning-informed Trajectory Prediction for Autonomous Driving
- Decision-Making for Automated Vehicles Using a Hierarchical Behavior-Based Arbitration Scheme
- Mind the gaps: Assuring the safety of autonomous systems from an engineering, ethical, and legal perspective
- Learning by Cheating
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Urban Driving with Conditional Imitation Learning
- End-to-End Model-Free Reinforcement Learning for Urban Driving using Implicit Affordances
- Looking at the right stuff: Guided semantic-gaze for autonomous driving
- CoverNet: Multimodal Behavior Prediction using Trajectory Sets
- Multiple Futures Prediction
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- A survey of deep learning techniques for autonomous driving
- Jointly Learnable Behavior and Trajectory Planning for Self-Driving Vehicles
- INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps
- A Survey of Autonomous Driving: Common Practices and Emerging Technologies
- Argoverse: 3D Tracking and Forecasting with Rich Maps
- End-to-end Interpretable Neural Motion Planner
- Exploring the Limitations of Behavior Cloning for Autonomous Driving
- Large-Scale Long-Tailed Recognition in an Open World
- Deep Learning for Large-Scale Traffic-Sign Detection and Recognition
- nuScenes: A multimodal dataset for autonomous driving
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
- Capturing Object Detection Uncertainty in Multi-Layer Grid Maps
- Cyber Security Challenges and Solutions for V2X Communications: A Survey
- Multilingual Constituency Parsing with Self-Attention and Pre-Training
- IntentNet: Learning to Predict Intention from Raw Sensor Data
- Deep Imitative Models for Flexible Inference, Planning, and Control
- Attention-based Active Visual Search for Mobile Robots
- Conditional Neural Processes
- Learning Multi-Modal Self-Awareness Models for Autonomous Vehicles from Human Driving
- Driving Policy Transfer via Modularity and Abstraction
- The ApolloScape Open Dataset for Autonomous Driving and its Application
- Multimodal Probabilistic Model-Based Planning for Human-Robot Interaction
- End-to-end Driving via Conditional Imitation Learning
- Attention Is All You Need
- Dynamic Occupancy Grid Prediction for Urban Autonomous Driving: A Deep Learning Approach with Fully Automatic Labeling
- DESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents
- Deep Reinforcement Learning framework for Autonomous Driving
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- Concrete Problems in AI Safety
- Matching Networks for One Shot Learning
- End to End Learning for Self-Driving Cars
- A Survey of Motion Planning and Control Techniques for Self-driving Urban Vehicles
- Revisiting Active Perception
- Vision meets robotics: The KITTI dataset
- A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
- Situation Awareness: Proceed with Caution
- Toward a Theory of Situation Awareness in Dynamic Systems
Discussions
Related