2017/02/20 by Yevgeny Seldin, Gábor Lugosi, Seldin, Yevgeny +1 · 7 citations
Computer Science · Engineering · Mathematics · #Anomaly Detection Techniques and Applications #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1702.06103
openalex publication_date 2017/02/20 · openalex created_date 2017/03/16 · arxiv created 2017/05/09 · arxiv updated 2017/05/10 · openalex updated_date 2026/07/28
We present a new strategy for gap estimation in randomized algorithms for multiarmed bandits and combine it with the EXP3++ algorithm of Seldin and Slivkins (2014). In the stochastic regime the strategy reduces dependence of regret on a time horizon from (ln t)3 to (ln t)2 and eliminates an additive factor of order Δe1/Δ2, where Δ is the minimal gap of a problem instance. In the adversarial regime regret guarantee remains unchanged.