Learning in Bayesian Stackelberg Games With Unknown Follower’s Types

Matteo Bollini, Francesco Bacchiocchi, Samuel Coutts, Matteo Castiglioni, Alberto Marchesi
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:8970-8994, 2026.

Abstract

We study online learning in Bayesian Stackelberg games, where a leader repeatedly interacts with a follower whose unknown private type is independently drawn at each round from an unknown probability distribution. The goal is to design algorithms that minimize the leader’s regret with respect to always playing an optimal commitment computed with knowledge of the game. We consider, for the first time to the best of our knowledge, the most realistic case in which the leader does not know anything about follower’s types, i.e., the possible follower’s payoffs. This raises considerable additional challenges compared to the usually addressed case in which follower’s payoffs are known. First, we prove a strong negative result: no-regret is unattainable under action feedback, i.e., when the leader only observes the follower’s best response at the end of each round. Thus, we focus on the easier type feedback model, where the follower’s type is also revealed. In such a setting, we propose an algorithm that achieves a regret of $\widetilde{O}(\sqrt{T})$, ignoring the dependence on other parameters.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bollini26a, title = {Learning in {B}ayesian Stackelberg Games With Unknown Follower’s Types}, author = {Bollini, Matteo and Bacchiocchi, Francesco and Coutts, Samuel and Castiglioni, Matteo and Marchesi, Alberto}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {8970--8994}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bollini26a/bollini26a.pdf}, url = {https://proceedings.mlr.press/v306/bollini26a.html}, abstract = {We study online learning in Bayesian Stackelberg games, where a leader repeatedly interacts with a follower whose unknown private type is independently drawn at each round from an unknown probability distribution. The goal is to design algorithms that minimize the leader’s regret with respect to always playing an optimal commitment computed with knowledge of the game. We consider, for the first time to the best of our knowledge, the most realistic case in which the leader does not know anything about follower’s types, i.e., the possible follower’s payoffs. This raises considerable additional challenges compared to the usually addressed case in which follower’s payoffs are known. First, we prove a strong negative result: no-regret is unattainable under action feedback, i.e., when the leader only observes the follower’s best response at the end of each round. Thus, we focus on the easier type feedback model, where the follower’s type is also revealed. In such a setting, we propose an algorithm that achieves a regret of $\widetilde{O}(\sqrt{T})$, ignoring the dependence on other parameters.} }
Endnote
%0 Conference Paper %T Learning in Bayesian Stackelberg Games With Unknown Follower’s Types %A Matteo Bollini %A Francesco Bacchiocchi %A Samuel Coutts %A Matteo Castiglioni %A Alberto Marchesi %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bollini26a %I PMLR %P 8970--8994 %U https://proceedings.mlr.press/v306/bollini26a.html %V 306 %X We study online learning in Bayesian Stackelberg games, where a leader repeatedly interacts with a follower whose unknown private type is independently drawn at each round from an unknown probability distribution. The goal is to design algorithms that minimize the leader’s regret with respect to always playing an optimal commitment computed with knowledge of the game. We consider, for the first time to the best of our knowledge, the most realistic case in which the leader does not know anything about follower’s types, i.e., the possible follower’s payoffs. This raises considerable additional challenges compared to the usually addressed case in which follower’s payoffs are known. First, we prove a strong negative result: no-regret is unattainable under action feedback, i.e., when the leader only observes the follower’s best response at the end of each round. Thus, we focus on the easier type feedback model, where the follower’s type is also revealed. In such a setting, we propose an algorithm that achieves a regret of $\widetilde{O}(\sqrt{T})$, ignoring the dependence on other parameters.
APA
Bollini, M., Bacchiocchi, F., Coutts, S., Castiglioni, M. & Marchesi, A.. (2026). Learning in Bayesian Stackelberg Games With Unknown Follower’s Types. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:8970-8994 Available from https://proceedings.mlr.press/v306/bollini26a.html.

Related Material