On Convergence and Optimality of Best-Response Learning with Policy Types in Multiagent Systems

Stefano Albrecht, Subramanian Ramamoorthy
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:246-255, 2014.

Abstract

While many multiagent algorithms are designed for homogeneous systems (i.e. all agents are iden- tical), there are important applications which re- quire an agent to coordinate its actions without knowing a priori how the other agents behave. One method to make this problem feasible is to as- sume that the other agents draw their latent policy (or type) from a specific set, and that a domain ex- pert could provide a specification of this set, albeit only a partially correct one. Algorithms have been proposed by several researchers to compute poste- rior beliefs over such policy libraries, which can then be used to determine optimal actions. In this paper, we provide theoretical guidance on two cen- tral design parameters of this method: Firstly, it is important that the user choose a posterior which can learn the true distribution of latent types, as otherwise suboptimal actions may be chosen. We analyse convergence properties of two existing posterior formulations and propose a new poste- rior which can learn correlated distributions. Sec- ondly, since the types are provided by an expert, they may be inaccurate in the sense that they do not predict the agents’ observed actions. We pro- vide a novel characterisation of optimality which allows experts to use efficient model checking al- gorithms to verify optimality of types.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-albrecht14a, title = {On Convergence and Optimality of Best-Response Learning with Policy Types in Multiagent Systems}, author = {Albrecht, Stefano and Ramamoorthy, Subramanian}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {246--255}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/albrecht14a/albrecht14a.pdf}, url = {https://proceedings.mlr.press/r12/albrecht14a.html}, abstract = {While many multiagent algorithms are designed for homogeneous systems (i.e. all agents are iden- tical), there are important applications which re- quire an agent to coordinate its actions without knowing a priori how the other agents behave. One method to make this problem feasible is to as- sume that the other agents draw their latent policy (or type) from a specific set, and that a domain ex- pert could provide a specification of this set, albeit only a partially correct one. Algorithms have been proposed by several researchers to compute poste- rior beliefs over such policy libraries, which can then be used to determine optimal actions. In this paper, we provide theoretical guidance on two cen- tral design parameters of this method: Firstly, it is important that the user choose a posterior which can learn the true distribution of latent types, as otherwise suboptimal actions may be chosen. We analyse convergence properties of two existing posterior formulations and propose a new poste- rior which can learn correlated distributions. Sec- ondly, since the types are provided by an expert, they may be inaccurate in the sense that they do not predict the agents’ observed actions. We pro- vide a novel characterisation of optimality which allows experts to use efficient model checking al- gorithms to verify optimality of types.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T On Convergence and Optimality of Best-Response Learning with Policy Types in Multiagent Systems %A Stefano Albrecht %A Subramanian Ramamoorthy %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-albrecht14a %I PMLR %P 246--255 %U https://proceedings.mlr.press/r12/albrecht14a.html %V R12 %X While many multiagent algorithms are designed for homogeneous systems (i.e. all agents are iden- tical), there are important applications which re- quire an agent to coordinate its actions without knowing a priori how the other agents behave. One method to make this problem feasible is to as- sume that the other agents draw their latent policy (or type) from a specific set, and that a domain ex- pert could provide a specification of this set, albeit only a partially correct one. Algorithms have been proposed by several researchers to compute poste- rior beliefs over such policy libraries, which can then be used to determine optimal actions. In this paper, we provide theoretical guidance on two cen- tral design parameters of this method: Firstly, it is important that the user choose a posterior which can learn the true distribution of latent types, as otherwise suboptimal actions may be chosen. We analyse convergence properties of two existing posterior formulations and propose a new poste- rior which can learn correlated distributions. Sec- ondly, since the types are provided by an expert, they may be inaccurate in the sense that they do not predict the agents’ observed actions. We pro- vide a novel characterisation of optimality which allows experts to use efficient model checking al- gorithms to verify optimality of types. %Z Reissued by PMLR on 04 October 2026.
APA
Albrecht, S. & Ramamoorthy, S.. (2014). On Convergence and Optimality of Best-Response Learning with Policy Types in Multiagent Systems. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:246-255 Available from https://proceedings.mlr.press/r12/albrecht14a.html. Reissued by PMLR on 04 October 2026.

Related Material