Efficient Decentralized Learning of Generalized Quantal Response Equilibrium

Zehao Zhao, Apurv Shukla, Rahul Jain, Vijay G Subramanian
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:8182-8208, 2026.

Abstract

We study a solution concept for bounded rational agents in finite normal-form general-sum games called Generalized Quantal Response Equilibrium (GQRE) which generalizes a Quantal Response Equilibrium (McKelvey and Palfrey, 1995). In our setup, each player can individually maximize a smooth, regularized expected utility of the mixed profiles used, reflecting both bounded rationality that subsumes stochastic choice, and also individual choice of behaviors. After establishing existence under mild conditions, we present a computationally efficient no-regret decentralized learning algorithm that uses a smoothened version of the Frank–Wolfe algorithm coupled with a computationally efficient projection step. Our algorithm uses noisy gradient estimates via bandit-feedback from a simulation oracle that reports on repeated plays of the game. We analyze finite-time convergence properties of our algorithm under assumptions that ensure uniqueness of equilibrium, using a novel class of gap functions that generalize the {Nash} gap function. We end by demonstrating the effectiveness of our method on a set of complex general-sum games such as high-rank two-player games, large action two-player games, and known examples of difficult multi-player games.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-zhao26a, title = {Efficient Decentralized Learning of Generalized Quantal Response Equilibrium}, author = {Zhao, Zehao and Shukla, Apurv and Jain, Rahul and Subramanian, Vijay G}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {8182--8208}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/zhao26a/zhao26a.pdf}, url = {https://proceedings.mlr.press/v337/zhao26a.html}, abstract = {We study a solution concept for bounded rational agents in finite normal-form general-sum games called Generalized Quantal Response Equilibrium (GQRE) which generalizes a Quantal Response Equilibrium (McKelvey and Palfrey, 1995). In our setup, each player can individually maximize a smooth, regularized expected utility of the mixed profiles used, reflecting both bounded rationality that subsumes stochastic choice, and also individual choice of behaviors. After establishing existence under mild conditions, we present a computationally efficient no-regret decentralized learning algorithm that uses a smoothened version of the Frank–Wolfe algorithm coupled with a computationally efficient projection step. Our algorithm uses noisy gradient estimates via bandit-feedback from a simulation oracle that reports on repeated plays of the game. We analyze finite-time convergence properties of our algorithm under assumptions that ensure uniqueness of equilibrium, using a novel class of gap functions that generalize the {Nash} gap function. We end by demonstrating the effectiveness of our method on a set of complex general-sum games such as high-rank two-player games, large action two-player games, and known examples of difficult multi-player games.} }
Endnote
%0 Conference Paper %T Efficient Decentralized Learning of Generalized Quantal Response Equilibrium %A Zehao Zhao %A Apurv Shukla %A Rahul Jain %A Vijay G Subramanian %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-zhao26a %I PMLR %P 8182--8208 %U https://proceedings.mlr.press/v337/zhao26a.html %V 337 %X We study a solution concept for bounded rational agents in finite normal-form general-sum games called Generalized Quantal Response Equilibrium (GQRE) which generalizes a Quantal Response Equilibrium (McKelvey and Palfrey, 1995). In our setup, each player can individually maximize a smooth, regularized expected utility of the mixed profiles used, reflecting both bounded rationality that subsumes stochastic choice, and also individual choice of behaviors. After establishing existence under mild conditions, we present a computationally efficient no-regret decentralized learning algorithm that uses a smoothened version of the Frank–Wolfe algorithm coupled with a computationally efficient projection step. Our algorithm uses noisy gradient estimates via bandit-feedback from a simulation oracle that reports on repeated plays of the game. We analyze finite-time convergence properties of our algorithm under assumptions that ensure uniqueness of equilibrium, using a novel class of gap functions that generalize the {Nash} gap function. We end by demonstrating the effectiveness of our method on a set of complex general-sum games such as high-rank two-player games, large action two-player games, and known examples of difficult multi-player games.
APA
Zhao, Z., Shukla, A., Jain, R. & Subramanian, V.G.. (2026). Efficient Decentralized Learning of Generalized Quantal Response Equilibrium. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:8182-8208 Available from https://proceedings.mlr.press/v337/zhao26a.html.

Related Material