Bandit Social Learning with Exploration Episodes

Kiarash Banihashem, Natalie Collina, Aleksandrs Slivkins
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:6242-6280, 2026.

Abstract

We study a stylized social learning dynamics where self-interested agents collectively follow a simple multi-armed bandit protocol. Each agent controls an "episode": a short sequence of consecutive decisions. Motivating applications include users repeatedly interacting with an AI, or repeatedly shopping at a marketplace. While agents are incentivized to explore within their respective episodes, we show that the aggregate exploration fails: e.g., its Bayesian regret grows linearly over time. In fact, such failure is a (very) typical case, not just a worst-case scenario. This conclusion persists even if an agent’s per-episode utility is some fixed function of the per-round outcomes: e.g., $\min$ or $\max$, not just the sum. Thus, externally driven exploration is needed even when some amount of exploration happens organically.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-banihashem26a, title = {Bandit Social Learning with Exploration Episodes}, author = {Banihashem, Kiarash and Collina, Natalie and Slivkins, Aleksandrs}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {6242--6280}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/banihashem26a/banihashem26a.pdf}, url = {https://proceedings.mlr.press/v306/banihashem26a.html}, abstract = {We study a stylized social learning dynamics where self-interested agents collectively follow a simple multi-armed bandit protocol. Each agent controls an "episode": a short sequence of consecutive decisions. Motivating applications include users repeatedly interacting with an AI, or repeatedly shopping at a marketplace. While agents are incentivized to explore within their respective episodes, we show that the aggregate exploration fails: e.g., its Bayesian regret grows linearly over time. In fact, such failure is a (very) typical case, not just a worst-case scenario. This conclusion persists even if an agent’s per-episode utility is some fixed function of the per-round outcomes: e.g., $\min$ or $\max$, not just the sum. Thus, externally driven exploration is needed even when some amount of exploration happens organically.} }
Endnote
%0 Conference Paper %T Bandit Social Learning with Exploration Episodes %A Kiarash Banihashem %A Natalie Collina %A Aleksandrs Slivkins %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-banihashem26a %I PMLR %P 6242--6280 %U https://proceedings.mlr.press/v306/banihashem26a.html %V 306 %X We study a stylized social learning dynamics where self-interested agents collectively follow a simple multi-armed bandit protocol. Each agent controls an "episode": a short sequence of consecutive decisions. Motivating applications include users repeatedly interacting with an AI, or repeatedly shopping at a marketplace. While agents are incentivized to explore within their respective episodes, we show that the aggregate exploration fails: e.g., its Bayesian regret grows linearly over time. In fact, such failure is a (very) typical case, not just a worst-case scenario. This conclusion persists even if an agent’s per-episode utility is some fixed function of the per-round outcomes: e.g., $\min$ or $\max$, not just the sum. Thus, externally driven exploration is needed even when some amount of exploration happens organically.
APA
Banihashem, K., Collina, N. & Slivkins, A.. (2026). Bandit Social Learning with Exploration Episodes. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:6242-6280 Available from https://proceedings.mlr.press/v306/banihashem26a.html.

Related Material