vix.ing · top · new · best · stats · spec

Efficient Online Bandit Multiclass Learning with O(√(T)) Regret

2017/02/25 by Alina Beygelzimer, Francesco Orabona, Beygelzimer, Alina +3
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Optimization and Search Problems

paper · pdf · doi:10.48550/arxiv.1702.07958

openalex publication_date 2017/02/25 · openalex created_date 2017/03/16 · openalex updated_date 2026/07/28

Abstract

We present an efficient second-order algorithm with O(\frac1η√(T)) regret for the bandit online multiclass problem. The regret bound holds simultaneously with respect to a family of loss functions parameterized by η, for a range of η restricted by the norm of the competitor. The family of loss functions ranges from hinge loss (η=0) to squared hinge loss (η=1). This provides a solution to the open problem of (J. Abernethy and A. Rakhlin. An efficient bandit algorithm for √(T)-regret in online multiclass prediction? In COLT, 2009). We test our algorithm experimentally, showing that it also performs favorably against earlier algorithms.

Citations

Related