2024/09/13 by Rowan Swiers, Swiers, Rowan, Subash Prabanantham +3
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management
paper · pdf · doi:10.48550/arxiv.2409.09199
openalex publication_date 2024/09/13 · openalex created_date 2024/10/23 · openalex updated_date 2026/07/28
Multi-armed Bandits (MABs) are increasingly employed in online platforms and e-commerce to optimize decision making for personalized user experiences. In this work, we focus on the Contextual Bandit problem with linear rewards, under conditions of sparsity and batched data. We address the challenge of fairness by excluding irrelevant features from decision-making processes using a novel algorithm, Online Batched Sequential Inclusion (OBSI), which sequentially includes features as confidence in their impact on the reward increases. Our experiments on synthetic data show the superior performance of OBSI compared to other algorithms in terms of regret, relevance of features used, and compute.