2020/07/10 by Ronan Fruit, Fruit, Ronan, Matteo Pirotta +3 · 1 citation
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2007.05456
openalex publication_date 2020/07/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider the problem of exploration-exploitation in communicating Markov Decision Processes. We provide an analysis of UCRL2 with Empirical Bernstein inequalities (UCRL2B). For any MDP with S states, A actions, Γ≤ S next states and diameter D, the regret of UCRL2B is bounded as \widetildeO(√(DΓS A T)).