vix.ing · top · new · best · stats

Online Semi-Supervised Learning in Contextual Bandits with Episodic Reward

2020/09/17 by Baihan Lin, Lin, Baihan
Computer Science · Decision Sciences · Engineering · Mathematics · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management #cs.AI #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2009.08457

Proceeding of AJCAI 2020. This article supersedes our work arXiv:1802.00981 on contextual bandits in nonstationary setting, introduces a new problem setting with episodically revealed reward, and provides a novel solution by propagating pseudo-feedbacks to un-rewarded cases from self-supervision. Also check out our speaker diarization application of this algorithm at arXiv:2006.04376

openalex publication_date 2020/09/17 · arxiv created 2020/10/25 · arxiv updated 2020/10/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We considered a novel practical problem of online learning with episodically revealed rewards, motivated by several real-world applications, where the contexts are nonstationary over different episodes and the reward feedbacks are not always available to the decision making agents. For this online semi-supervised learning setting, we introduced Background Episodic Reward LinUCB (BerlinUCB), a solution that easily incorporates clustering as a self-supervision module to provide useful side information when rewards are not observed. Our experiments on a variety of datasets, both in stationary and nonstationary environments of six different scenarios, demonstrated clear advantages of the proposed approach over the standard contextual bandit. Lastly, we introduced a relevant real-life example where this problem setting is especially useful.

Citations

Related