2018/04/02 by Anirudh Goyal, Goyal, Anirudh, Philémon Brakel +14 · 1 voice · 5 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1804.00379
Accepted at ICLR 2019
openalex publication_date 2018/04/02 · arxiv published 2018/04/02 · arxiv created 2019/01/28 · arxiv updated 2019/01/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In many environments only a tiny subset of all states yield high reward. In these cases, few of the interactions with the environment provide a relevant learning signal. Hence, we may want to preferentially train on those high-reward states and the probable trajectories leading to them. To this end, we advocate for the use of a backtracking model that predicts the preceding states that terminate at a given high-reward state. We can train a model which, starting from a high value state (or one that is estimated to have high value), predicts and sample for which the (state, action)-tuples may have led to that high value state. These traces of (state, action) pairs, which we refer to as Recall Traces, sampled from this backtracking model starting from a high value state, are informative as they terminate in good states, and hence we can use these traces to improve a policy. We provide a variational interpretation for this idea and a practical algorithm in which the backtracking model samples from an approximate posterior distribution over trajectories which lead to large rewards. Our method improves the sample efficiency of both on- and off-policy RL algorithms across several environments and tasks.