Meet Me at the Arm: The Cooperative Multi Armed Bandits Problem with Shareable Arms

Xinyi Hu, Aldo Pacchiano
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1090-1098, 2026.

Abstract

We study the decentralized multi-player multi-armed bandits (MMAB) problem under a no-sensing setting, where each player receives only their own reward and obtains no information about collisions. Each arm has an unknown capacity, and if the number of players pulling an arm exceeds its capacity, all players involved receive zero reward. This setting generalizes the classical unit-capacity model and introduces new challenges in coordination and capacity discovery under severe feedback limitations. We propose A-CAPELLA (Algorithm for Capacity-Aware Parallel Elimination for Learning and Allocation), a decentralized learning algorithm that achieves logarithmic regret in this generalized regime via protocol-driven coordination.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-hu26a, title = { Meet Me at the Arm: The Cooperative Multi Armed Bandits Problem with Shareable Arms }, author = {Hu, Xinyi and Pacchiano, Aldo}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1090--1098}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/hu26a/hu26a.pdf}, url = {https://proceedings.mlr.press/v300/hu26a.html}, abstract = { We study the decentralized multi-player multi-armed bandits (MMAB) problem under a no-sensing setting, where each player receives only their own reward and obtains no information about collisions. Each arm has an unknown capacity, and if the number of players pulling an arm exceeds its capacity, all players involved receive zero reward. This setting generalizes the classical unit-capacity model and introduces new challenges in coordination and capacity discovery under severe feedback limitations. We propose A-CAPELLA (Algorithm for Capacity-Aware Parallel Elimination for Learning and Allocation), a decentralized learning algorithm that achieves logarithmic regret in this generalized regime via protocol-driven coordination. } }
Endnote
%0 Conference Paper %T Meet Me at the Arm: The Cooperative Multi Armed Bandits Problem with Shareable Arms %A Xinyi Hu %A Aldo Pacchiano %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-hu26a %I PMLR %P 1090--1098 %U https://proceedings.mlr.press/v300/hu26a.html %V 300 %X We study the decentralized multi-player multi-armed bandits (MMAB) problem under a no-sensing setting, where each player receives only their own reward and obtains no information about collisions. Each arm has an unknown capacity, and if the number of players pulling an arm exceeds its capacity, all players involved receive zero reward. This setting generalizes the classical unit-capacity model and introduces new challenges in coordination and capacity discovery under severe feedback limitations. We propose A-CAPELLA (Algorithm for Capacity-Aware Parallel Elimination for Learning and Allocation), a decentralized learning algorithm that achieves logarithmic regret in this generalized regime via protocol-driven coordination.
APA
Hu, X. & Pacchiano, A.. (2026). Meet Me at the Arm: The Cooperative Multi Armed Bandits Problem with Shareable Arms . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1090-1098 Available from https://proceedings.mlr.press/v300/hu26a.html.

Related Material