2020/08/14 by Cristina Pinneri, Shambhuraj Sawant, Pinneri, Cristina +12 · 18 citations
Computer Science · Engineering · Mathematics · #Advanced Control Systems Optimization #Reinforcement Learning in Robotics #Robotic Path Planning Algorithms #cs.LG #cs.RO #stat.ML
paper · pdf · doi:10.48550/arxiv.2008.06389
arxiv created 2020/08/14 · arxiv updated 2020/08/17
Trajectory optimizers for model-based reinforcement learning, such as the Cross-Entropy Method (CEM), can yield compelling results even in high-dimensional control tasks and sparse-reward environments. However, their sampling inefficiency prevents them from being used for real-time planning and control. We propose an improved version of the CEM algorithm for fast planning, with novel additions including temporally-correlated actions and memory, requiring 2.7-22x less samples and yielding a performance increase of 1.2-10x in high-dimensional control problems.