2021/12/13 by Sergey Garbar, Garbar, Sergey
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Distributed Sensor Networks and Detection Algorithms #FOS: Computer and information sciences #FOS: Mathematics #Forecasting Techniques and Applications #Machine Learning (cs.LG) #Optimization and Control (math.OC) #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2112.06423
openalex publication_date 2021/12/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider the upper confidence bound strategy for Gaussian multi-armed bandits with known control horizon sizes N and build its limiting description with a system of stochastic differential equations and ordinary differential equations. Rewards for the arms are assumed to have unknown expected values and known variances. A set of Monte-Carlo simulations was performed for the case of close distributions of rewards, when mean rewards differ by the magnitude of order N-1/2, as it yields the highest normalized regret, to verify the validity of the obtained description. The minimal size of the control horizon when the normalized regret is not noticeably larger than maximum possible was estimated.