2020/08/27 by Fei Ye, Pin Wang, Ye, Fei +5
Engineering · #Autonomous Vehicle Technology and Safety #Traffic control and management #Vehicle emissions and performance
paper · pdf · doi:10.48550/arxiv.2008.12451
Recent advances in supervised learning and reinforcement learning have\nprovided new opportunities to apply related methodologies to automated driving.\nHowever, there are still challenges to achieve automated driving maneuvers in\ndynamically changing environments. Supervised learning algorithms such as\nimitation learning can generalize to new environments by training on a large\namount of labeled data, however, it can be often impractical or\ncost-prohibitive to obtain sufficient data for each new environment. Although\nreinforcement learning methods can mitigate this data-dependency issue by\ntraining the agent in a trial-and-error way, they still need to re-train\npolicies from scratch when adapting to new environments. In this paper, we thus\npropose a meta reinforcement learning (MRL) method to improve the agent's\ngeneralization capabilities to make automated lane-changing maneuvers at\ndifferent traffic environments, which are formulated as different traffic\ncongestion levels. Specifically, we train the model at light to moderate\ntraffic densities and test it at a new heavy traffic density condition. We use\nboth collision rate and success rate to quantify the safety and effectiveness\nof the proposed model. A benchmark model is developed based on a pretraining\nmethod, which uses the same network structure and training tasks as our\nproposed model for fair comparison. The simulation results shows that the\nproposed method achieves an overall success rate up to 20% higher than the\nbenchmark model when it is generalized to the new environment of heavy traffic\ndensity. The collision rate is also reduced by up to 18% than the benchmark\nmodel. Finally, the proposed model shows more stable and efficient\ngeneralization capabilities adapting to the new environment, and it can achieve\n100% successful rate and 0% collision rate with only a few steps of gradient\nupdates.\n