What Preferences Can—and Cannot—Predict in Multi-Agent Online Learning

Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:22-68, 2026.

Abstract

We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game—its preference graph—determine the outcomes of no-regret learning dynamics—such as follow-the-regularized-leader (FTRL). In one direction, we show that the skeleton of every dynamically stable set (i.e. the set of pure profiles it contains) must also be preferentially stable, that is, it must be closed under profitable deviations. We then ask the converse question: when do preferences determine the long-run behavior of the players’ learning dynamics? We begin by showing that preferences characterize asymptotic stability in the case of subgames—i.e. subsets of pure profiles obtained by restricting players’ action sets. Beyond this case however, the equivalence between dynamic and preferential stability collapses: concretely, we construct a three-player game with a preferentially stable set whose span is dynamically unstable, showing in this way that preferences do not suffice as a criterion of dynamic stability. We then bridge this gap via the notion of resilience under aggregate deviations, an easy-to-check payoff-based condition that guarantees asymptotic stability of arbitrary spans of pure strategies.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-abbadi26a, title = {What Preferences Can—and Cannot—Predict in Multi-Agent Online Learning}, author = {Abbadi, Omar and Laraki, Rida and Mertikopoulos, Panayotis}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {22--68}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/abbadi26a/abbadi26a.pdf}, url = {https://proceedings.mlr.press/v306/abbadi26a.html}, abstract = {We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game—its preference graph—determine the outcomes of no-regret learning dynamics—such as follow-the-regularized-leader (FTRL). In one direction, we show that the skeleton of every dynamically stable set (i.e. the set of pure profiles it contains) must also be preferentially stable, that is, it must be closed under profitable deviations. We then ask the converse question: when do preferences determine the long-run behavior of the players’ learning dynamics? We begin by showing that preferences characterize asymptotic stability in the case of subgames—i.e. subsets of pure profiles obtained by restricting players’ action sets. Beyond this case however, the equivalence between dynamic and preferential stability collapses: concretely, we construct a three-player game with a preferentially stable set whose span is dynamically unstable, showing in this way that preferences do not suffice as a criterion of dynamic stability. We then bridge this gap via the notion of resilience under aggregate deviations, an easy-to-check payoff-based condition that guarantees asymptotic stability of arbitrary spans of pure strategies.} }
Endnote
%0 Conference Paper %T What Preferences Can—and Cannot—Predict in Multi-Agent Online Learning %A Omar Abbadi %A Rida Laraki %A Panayotis Mertikopoulos %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-abbadi26a %I PMLR %P 22--68 %U https://proceedings.mlr.press/v306/abbadi26a.html %V 306 %X We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game—its preference graph—determine the outcomes of no-regret learning dynamics—such as follow-the-regularized-leader (FTRL). In one direction, we show that the skeleton of every dynamically stable set (i.e. the set of pure profiles it contains) must also be preferentially stable, that is, it must be closed under profitable deviations. We then ask the converse question: when do preferences determine the long-run behavior of the players’ learning dynamics? We begin by showing that preferences characterize asymptotic stability in the case of subgames—i.e. subsets of pure profiles obtained by restricting players’ action sets. Beyond this case however, the equivalence between dynamic and preferential stability collapses: concretely, we construct a three-player game with a preferentially stable set whose span is dynamically unstable, showing in this way that preferences do not suffice as a criterion of dynamic stability. We then bridge this gap via the notion of resilience under aggregate deviations, an easy-to-check payoff-based condition that guarantees asymptotic stability of arbitrary spans of pure strategies.
APA
Abbadi, O., Laraki, R. & Mertikopoulos, P.. (2026). What Preferences Can—and Cannot—Predict in Multi-Agent Online Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:22-68 Available from https://proceedings.mlr.press/v306/abbadi26a.html.

Related Material