2019/09/15 by Zac Wellmer, Wellmer, Zac, James T. Kwok +1
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1909.07373
openalex publication_date 2019/09/15 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28
This paper proposes a novel deep reinforcement learning architecture that was\ninspired by previous tree structured architectures which were only useable in\ndiscrete action spaces. Policy Prediction Network offers a way to improve\nsample complexity and performance on continuous control problems in exchange\nfor extra computation at training time but at no cost in computation at rollout\ntime. Our approach integrates a mix between model-free and model-based\nreinforcement learning. Policy Prediction Network is the first to introduce\nimplicit model-based learning to Policy Gradient algorithms for continuous\naction space and is made possible via the empirically justified clipping\nscheme. Our experiments are focused on the MuJoCo environments so that they can\nbe compared with similar work done in this area.\n