vix.ing · top · new · best · stats · spec

OptionGAN: Learning Joint Reward-Policy Options using Generative\n Adversarial Inverse Reinforcement Learning

2017/09/19 by Peter Henderson, Henderson, Peter, Wei-Di Chang +9 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.1709.06683

openalex publication_date 2017/09/19 · openalex created_date 2022/10/04 · openalex updated_date 2026/07/28

Abstract

Reinforcement learning has shown promise in learning policies that can solve\ncomplex problems. However, manually specifying a good reward function can be\ndifficult, especially for intricate tasks. Inverse reinforcement learning\noffers a useful paradigm to learn the underlying reward function directly from\nexpert demonstrations. Yet in reality, the corpus of demonstrations may contain\ntrajectories arising from a diverse set of underlying reward functions rather\nthan a single one. Thus, in inverse reinforcement learning, it is useful to\nconsider such a decomposition. The options framework in reinforcement learning\nis specifically designed to decompose policies in a similar light. We therefore\nextend the options framework and propose a method to simultaneously recover\nreward options in addition to policy options. We leverage adversarial methods\nto learn joint reward-policy options using only observed expert states. We show\nthat this approach works well in both simple and complex continuous control\ntasks and shows significant performance increases in one-shot transfer\nlearning.\n

Cited by

Related