2019/06/12 by Botao Hao, Yasin Abbasi-Yadkori, Hao, Botao +5 · 8 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Distributed Sensor Networks and Detection Algorithms #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1906.05247
Accepted by NeurIPS 2019
openalex publication_date 2019/06/12 · arxiv created 2019/10/31 · arxiv updated 2019/11/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Upper Confidence Bound (UCB) method is arguably the most celebrated one used in online decision making with partial information feedback. Existing techniques for constructing confidence bounds are typically built upon various concentration inequalities, which thus lead to over-exploration. In this paper, we propose a non-parametric and data-dependent UCB algorithm based on the multiplier bootstrap. To improve its finite sample performance, we further incorporate second-order correction into the above construction. In theory, we derive both problem-dependent and problem-independent regret bounds for multi-armed bandits under a much weaker tail assumption than the standard sub-Gaussianity. Numerical results demonstrate significant regret reductions by our method, in comparison with several baselines in a range of multi-armed and linear bandit problems.