2021/10/26 by Junsu Kim, Kim, Junsu, Younggyo Seo +3 · 2 citations
Computer Science · Engineering · #Adaptive Dynamic Programming Control #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Fuel Cells and Related Materials #Machine Learning (cs.LG) #Muscle activation and electromyography studies #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2110.13625
openalex publication_date 2021/10/26 · openalex created_date 2021/11/22 · openalex updated_date 2026/07/28
Goal-conditioned hierarchical reinforcement learning (HRL) has shown\npromising results for solving complex and long-horizon RL tasks. However, the\naction space of high-level policy in the goal-conditioned HRL is often large,\nso it results in poor exploration, leading to inefficiency in training. In this\npaper, we present HIerarchical reinforcement learning Guided by Landmarks\n(HIGL), a novel framework for training a high-level policy with a reduced\naction space guided by landmarks, i.e., promising states to explore. The key\ncomponent of HIGL is twofold: (a) sampling landmarks that are informative for\nexploration and (b) encouraging the high-level policy to generate a subgoal\ntowards a selected landmark. For (a), we consider two criteria: coverage of the\nentire visited state space (i.e., dispersion of states) and novelty of states\n(i.e., prediction error of a state). For (b), we select a landmark as the very\nfirst landmark in the shortest path in a graph whose nodes are landmarks. Our\nexperiments demonstrate that our framework outperforms prior-arts across a\nvariety of control tasks, thanks to efficient exploration guided by landmarks.\n