2021/10/06 by Ting-Han Fan, Fan, Ting-Han, Peter J. Ramadge +1
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Age of Information Optimization #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2110.02421
arxiv created 2021/10/06 · openalex publication_date 2021/10/06 · arxiv updated 2021/10/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Off-policy Actor-Critic algorithms have demonstrated phenomenal experimental performance but still require better explanations. To this end, we show its policy evaluation error on the distribution of transitions decomposes into: a Bellman error, a bias from policy mismatch, and a variance term from sampling. By comparing the magnitude of bias and variance, we explain the success of the Emphasizing Recent Experience sampling and 1/age weighted sampling. Both sampling strategies yield smaller bias and variance and are hence preferable to uniform sampling.