vix.ing · top · new · best · stats · spec

\cal Q \cal D-Learning: A Collaborative Distributed Strategy for Multi-Agent Reinforcement Learning Through \rm Consensus + \rm Innovations

2012/05/31 by Soummya Kar, José M. F. Moura, Jose' M. F. Moura +1 · 1 citation
Computer Science · Mathematics · #Age of Information Optimization #Algorithm #Artificial intelligence #Bellman equation #Computer science #Distributed Control Multi-Agent Systems #Machine learning #Markov chain #Markov decision process #Markov process #Mathematical optimization #Mathematics #Reinforcement Learning in Robotics #Reinforcement learning #State (computer science) #Statistics #cs.LG #cs.MA #math.OC #math.PR #stat.ML

paper · pdf · doi:10.1109/tsp.2013.2241057

Submitted to the IEEE Transactions on Signal Processing, 33 pages

arxiv created 2012/10/25 · openalex publication_date 2013/01/18 · arxiv updated 2015/06/04 · openalex created_date 2016/06/24 · openalex updated_date 2026/08/05

Abstract

The paper developsQ D-learning, a distributed version of reinforcementQ-learning, for multi-agent Markov decision processes (MDPs); the agents have no prior information on the global state transition and on the local agent cost statistics. The network agents minimize a network-averaged infinite horizon discounted cost, by local processing and by collaborating through mutual information exchange over a sparse (possibly stochastic) communication network. The agents respond differently (depending on their instantaneous one-stage random costs) to a global controlled state and the control actions of a remote controller. When each agent is aware only of its local online cost data and the inter-agent communication network is weakly connected, we prove thatQ D-learning, a consensus + innovations algorithm with mixed time-scale stochastic dynamics, converges asymptotically almost surely to the desired value function and to the optimal stationary control policy at each network agent.

Citations

Cited by

Related