2018/09/18 by Izumi Karino, Karino, Izumi, Kazutoshi Tanaka +5
Computer Science · Physics and Astronomy · #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1809.06570
openalex publication_date 2018/09/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper proposes an exploration method for deep reinforcement learning\nbased on parameter space noise. Recent studies have experimentally shown that\nparameter space noise results in better exploration than the commonly used\naction space noise. Previous methods devised a way to update the diagonal\ncovariance matrix of a noise distribution and did not consider the direction of\nthe noise vector and its correlation. In addition, fast updates of the noise\ndistribution are required to facilitate policy learning. We propose a method\nthat deforms the noise distribution according to the accumulated returns and\nthe noises that have led to the returns. Moreover, this method switches\nisotropic exploration and directional exploration in parameter space with\nregard to obtained rewards. We validate our exploration strategy in the OpenAI\nGym continuous environments and modified environments with sparse rewards. The\nproposed method achieves results that are competitive with a previous method at\nbaseline tasks. Moreover, our approach exhibits better performance in sparse\nreward environments by exploration with the switching strategy.\n