2019/02/06 by Erhan Bayraktar, Bayraktar, Erhan, Ibrahim Ekren +3 · 1 citation
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning and Algorithms #Optimization and Control (math.OC) #Probability (math.PR) #Reinforcement Learning in Robotics #cs.LG #math.OC #math.PR
paper · pdf · doi:10.48550/arxiv.1902.02368
To appear in Annals of Applied Probabilty
openalex publication_date 2019/02/06 · openalex created_date 2019/02/21 · arxiv created 2020/03/19 · arxiv updated 2020/03/23 · openalex updated_date 2026/07/28
For the problem of prediction with expert advice in the adversarial setting with geometric stopping, we compute the exact leading order expansion for the long time behavior of the value function. Then, we use this expansion to prove that as conjectured in Gravin et al. [12], the comb strategies are indeed asymptotically optimal for the adversary in the case of 4 experts.