2025/02/20 by Alexander Terenin, Jeffrey Negrea, Terenin, Alexander +1 · 1 voice
#cs.LG #cs.GT #math.ST #stat.ML
paper · pdf · doi:10.48550/arxiv.2502.14790
We develop a form Thompson sampling for online learning under full feedback - also known as prediction with expert advice - where the learner's prior is defined over the space of an adversary's future actions, rather than the space of experts. We show regret decomposes into regret the learner expected a priori, plus a prior-robustness-type term we call excess regret. In the classical finite-expert setting, this recovers optimal rates. As an initial step towards practical online learning in settings with a potentially-uncountably-infinite number of experts, we show that Thompson sampling over the d-dimensional unit cube, using a certain Gaussian process prior widely-used in the Bayesian optimization literature, has a O(β√(Tdlog(1+√(d)\fracλβ))) rate against a β-bounded λ-Lipschitz adversary.