2015/08/14 by Assaf Hallak, Hallak, Assaf, Aviv Tamar +3
Computer Science · Engineering · #Advanced Control Systems Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1508.03411
openalex publication_date 2015/08/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recently, \citetSuttonMW15 introduced the emphatic temporal differences (ETD) algorithm for off-policy evaluation in Markov decision processes. In this short note, we show that the projected fixed-point equation that underlies ETD involves a contraction operator, with a √γ-contraction modulus (where γ is the discount factor). This allows us to provide error bounds on the approximation error of ETD. To our knowledge, these are the first error bounds for an off-policy evaluation algorithm under general target and behavior policies.