vix.ing · top · new · best · stats · spec

Black-box Off-policy Estimation for Infinite-Horizon Reinforcement\n Learning

2020/03/24 by Alireza Mousavi, Mousavi, Ali, Lihong Li +6 · 1 citation
Computer Science · Decision Sciences · #Adaptive Dynamic Programming Control #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2003.11126

openalex publication_date 2020/03/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Off-policy estimation for long-horizon problems is important in many\nreal-life applications such as healthcare and robotics, where high-fidelity\nsimulators may not be available and on-policy evaluation is expensive or\nimpossible. Recently, citeliu18breaking proposed an approach that avoids the\n\curse of horizon suffered by typical importance-sampling-based methods.\nWhile showing promising results, this approach is limited in practice as it\nrequires data be drawn from the \stationary distribution of a\n\known behavior policy. In this work, we propose a novel approach that\neliminates such limitations. In particular, we formulate the problem as solving\nfor the fixed point of a certain operator. Using tools from Reproducing Kernel\nHilbert Spaces (RKHSs), we develop a new estimator that computes importance\nratios of stationary distributions, without knowledge of how the off-policy\ndata are collected. We analyze its asymptotic consistency and finite-sample\ngeneralization. Experiments on benchmarks verify the effectiveness of our\napproach.\n

Citations

Cited by

Related