2016/06/02 by Peter Landgren, Landgren, Peter, Vaibhav Srivastava +3 · 3 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Auction Theory and Applications #FOS: Computer and information sciences #FOS: Electrical engineering #FOS: Mathematics #Machine Learning (cs.LG) #Mobile Crowdsensing and Crowdsourcing #Optimization and Control (math.OC) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1606.00911
openalex publication_date 2016/06/02 · openalex created_date 2022/10/02 · openalex updated_date 2026/07/28
We study distributed cooperative decision-making under the explore-exploit\ntradeoff in the multiarmed bandit (MAB) problem. We extend the state-of-the-art\nfrequentist and Bayesian algorithms for single-agent MAB problems to\ncooperative distributed algorithms for multi-agent MAB problems in which agents\ncommunicate according to a fixed network graph. We rely on a running consensus\nalgorithm for each agent's estimation of mean rewards from its own rewards and\nthe estimated rewards of its neighbors. We prove the performance of these\nalgorithms and show that they asymptotically recover the performance of a\ncentralized agent. Further, we rigorously characterize the influence of the\ncommunication graph structure on the decision-making performance of the group.\n