2023/02/09 by Dan Garber, Garber, Dan, Ben Kretzu +1 · 1 citation
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Optimization and Control (math.OC) #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2302.04859
openalex publication_date 2023/02/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider the setting of online convex optimization (OCO) with exp-concave losses. The best regret bound known for this setting is O(nlogT), where n is the dimension and T is the number of prediction rounds (treating all other quantities as constants and assuming T is sufficiently large), and is attainable via the well-known Online Newton Step algorithm (ONS). However, ONS requires on each iteration to compute a projection (according to some matrix-induced norm) onto the feasible convex set, which is often computationally prohibitive in high-dimensional settings and when the feasible set admits a non-trivial structure. In this work we consider projection-free online algorithms for exp-concave and smooth losses, where by projection-free we refer to algorithms that rely only on the availability of a linear optimization oracle (LOO) for the feasible set, which in many applications of interest admits much more efficient implementations than a projection oracle. We present an LOO-based ONS-style algorithm, which using overall O(T) calls to a LOO, guarantees in worst case regret bounded by \widetildeO(n2/3T2/3) (ignoring all quantities except for n,T). However, our algorithm is most interesting in an important and plausible low-dimensional data scenario: if the gradients (approximately) span a subspace of dimension at most ρ, ρ<< n, the regret bound improves to \widetildeO(ρ2/3T2/3), and by applying standard deterministic sketching techniques, both the space and average additional per-iteration runtime requirements are only O(ρn) (instead of O(n2)). This improves upon recently proposed LOO-based algorithms for OCO which, while having the same state-of-the-art dependence on the horizon T, suffer from regret/oracle complexity that scales with √(n) or worse.