2021/03/16 by Pranav Agarwal, Agarwal, Pranav, Pierre de Beaucorps +3
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Reinforcement Learning in Robotics #Robotics (cs.RO) #Transportation and Mobility Innovations
paper · pdf · doi:10.48550/arxiv.2103.09189
openalex publication_date 2021/03/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Deep reinforcement Learning for end-to-end driving is limited by the need of complex reward engineering. Sparse rewards can circumvent this challenge but suffers from long training time and leads to sub-optimal policy. In this work, we explore full-control driving with only goal-constrained sparse reward and propose a curriculum learning approach for end-to-end driving using only navigation view maps that benefit from small virtual-to-real domain gap. To address the complexity of multiple driving policies, we learn concurrent individual policies selected at inference by a navigation system. We demonstrate the ability of our proposal to generalize on unseen road layout, and to drive significantly longer than in the training.