2017/08/22 by Denis Steckelmacher, Steckelmacher, Denis, Diederik M. Roijers +9 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1708.06551
openalex publication_date 2017/08/22 · openalex created_date 2022/10/04 · openalex updated_date 2026/07/28
Many real-world reinforcement learning problems have a hierarchical nature,\nand often exhibit some degree of partial observability. While hierarchy and\npartial observability are usually tackled separately (for instance by combining\nrecurrent neural networks and options), we show that addressing both problems\nsimultaneously is simpler and more efficient in many cases. More specifically,\nwe make the initiation set of options conditional on the previously-executed\noption, and show that options with such Option-Observation Initiation Sets\n(OOIs) are at least as expressive as Finite State Controllers (FSCs), a\nstate-of-the-art approach for learning in POMDPs. OOIs are easy to design based\non an intuitive description of the task, lead to explainable policies and keep\nthe top-level and option policies memoryless. Our experiments show that OOIs\nallow agents to learn optimal policies in challenging POMDPs, while being much\nmore sample-efficient than a recurrent neural network over options.\n