2018/09/27 by Yunhao Tang, Tang, Yunhao, Shipra Agrawal +1 · 6 citations
Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Optimization and Search Problems #Reinforcement Learning in Robotics #Stochastic Gradient Optimization Techniques #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1809.10326
openalex publication_date 2018/09/27 · arxiv created 2019/02/01 · arxiv updated 2019/02/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose to improve trust region policy search with normalizing flows policy. We illustrate that when the trust region is constructed by KL divergence constraints, normalizing flows policy generates samples far from the 'center' of the previous policy iterate, which potentially enables better exploration and helps avoid bad local optima. Through extensive comparisons, we show that the normalizing flows policy significantly improves upon baseline architectures especially on high-dimensional tasks with complex dynamics.