vix.ing · top · new · best · stats · spec

Learning by Repetition: Stochastic Multi-armed Bandits under Priming\n Effect

2020/06/18 by Priyank Agrawal, Agrawal, Priyank, Theja Tulabandhula +1
Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management

paper · pdf · doi:10.48550/arxiv.2006.10356

openalex publication_date 2020/06/18 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

We study the effect of persistence of engagement on learning in a stochastic\nmulti-armed bandit setting. In advertising and recommendation systems,\nrepetition effect includes a wear-in period, where the user's propensity to\nreward the platform via a click or purchase depends on how frequently they see\nthe recommendation in the recent past. It also includes a counteracting\nwear-out period, where the user's propensity to respond positively is dampened\nif the recommendation was shown too many times recently. Priming effect can be\nnaturally modelled as a temporal constraint on the strategy space, since the\nreward for the current action depends on historical actions taken by the\nplatform. We provide novel algorithms that achieves sublinear regret in time\nand the relevant wear-in/wear-out parameters. The effect of priming on the\nregret upper bound is also additive, and we get back a guarantee that matches\npopular algorithms such as the UCB1 and Thompson sampling when there is no\npriming effect. Our work complements recent work on modeling time varying\nrewards, delays and corruptions in bandits, and extends the usage of rich\nbehavior models in sequential decision making settings.\n

Citations

Related