2021/02/18 by Nikos Karampatziakis, Karampatziakis, Nikos, Paul Mineiro +3 · 1 citation
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Age of Information Optimization #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2102.09540
openalex publication_date 2021/02/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We develop confidence bounds that hold uniformly over time for off-policy evaluation in the contextual bandit setting. These confidence sequences are based on recent ideas from martingale analysis and are non-asymptotic, non-parametric, and valid at arbitrary stopping times. We provide algorithms for computing these confidence sequences that strike a good balance between computational and statistical efficiency. We empirically demonstrate the tightness of our approach in terms of failure probability and width and apply it to the "gated deployment" problem of safely upgrading a production contextual bandit system.