vix.ing · top · new · best · stats · spec

On Convergence and Optimality of Best-Response Learning with Policy\n Types in Multiagent Systems

2019/07/15 by Stefano V. Albrecht, Subramanian Ramamoorthy, Albrecht, Stefano V. +1 · 1 citation
Computer Science · #Reinforcement Learning in Robotics #Machine Learning and Algorithms #Formal Methods in Verification

paper · pdf · doi:10.48550/arxiv.1907.06995

Abstract

While many multiagent algorithms are designed for homogeneous systems (i.e.\nall agents are identical), there are important applications which require an\nagent to coordinate its actions without knowing a priori how the other agents\nbehave. One method to make this problem feasible is to assume that the other\nagents draw their latent policy (or type) from a specific set, and that a\ndomain expert could provide a specification of this set, albeit only a\npartially correct one. Algorithms have been proposed by several researchers to\ncompute posterior beliefs over such policy libraries, which can then be used to\ndetermine optimal actions. In this paper, we provide theoretical guidance on\ntwo central design parameters of this method: Firstly, it is important that the\nuser choose a posterior which can learn the true distribution of latent types,\nas otherwise suboptimal actions may be chosen. We analyse convergence\nproperties of two existing posterior formulations and propose a new posterior\nwhich can learn correlated distributions. Secondly, since the types are\nprovided by an expert, they may be inaccurate in the sense that they do not\npredict the agents' observed actions. We provide a novel characterisation of\noptimality which allows experts to use efficient model checking algorithms to\nverify optimality of types.\n

Cited by

Related