vix.ing · top · new · best · stats · spec

The Journey is the Reward: Unsupervised Learning of Influential Trajectories

2019/05/22 by Jonathan Binas, Sherjil Ozair, Binas, Jonathan +3 · 1 voice
Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics #cs.AI #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1905.09334

openalex publication_date 2019/05/22 · arxiv published 2019/05/22 · arxiv updated 2019/05/22 · openalex created_date 2019/05/29 · openalex updated_date 2026/07/28

Abstract

Unsupervised exploration and representation learning become increasingly important when learning in diverse and sparse environments. The information-theoretic principle of empowerment formalizes an unsupervised exploration objective through an agent trying to maximize its influence on the future states of its environment. Previous approaches carry certain limitations in that they either do not employ closed-loop feedback or do not have an internal state. As a consequence, a privileged final state is taken as an influence measure, rather than the full trajectory. We provide a model-free method which takes into account the whole trajectory while still offering the benefits of option-based approaches. We successfully apply our approach to settings with large action spaces, where discovery of meaningful action sequences is particularly difficult.

Citations

Discussions

Related