Battle of Bandits

Aadirupa Saha, Aditya Gopalan
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:804-813, 2018.

Abstract

We introduce Battling-Bandits – an online learning framework where given a set of n arms, the learner needs to select a subset of k $\geq$2 arms in each round and subsequently observes a stochastic feedback indicating the winner of the round. This framework generalizes the stan- dard Dueling-Bandit framework which applies to several practical scenarios such as medical treatment preferences, recommender systems, search engine optimization etc., where it is eas- ier and more effective to collect feedback for multiple options simultaneously. We develop a novel class of pairwise-subset choice model, for modelling the subset-wise winner feedback and propose three algorithms - Battling-Doubler, Battling-MultiSBM and Battling-Duel: While the first two are designed for a special class of linear-link based choice models, the third one applies to a much general class of pairwise- subset choice models with Condorcet winner. We also analyzed their regret guarantees and show the optimality of Battling-Duel proving a matching regret lower bound of $\Omega$(n log T), which (perhaps surprisingly) shows that the flexibility of playing size-k subsets does not really help to gather information faster than the corresponding dueling case (k = 2), at least for the current subsetwise feedback choice model. The efficacy of our algorithms are demonstrated through extensive experimental evaluations on a variety of synthetic and real world datasets.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-saha18a, title = {Battle of Bandits}, author = {Saha, Aadirupa and Gopalan, Aditya}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {804--813}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/saha18a/saha18a.pdf}, url = {https://proceedings.mlr.press/r16/saha18a.html}, abstract = {We introduce Battling-Bandits – an online learning framework where given a set of n arms, the learner needs to select a subset of k $\geq$2 arms in each round and subsequently observes a stochastic feedback indicating the winner of the round. This framework generalizes the stan- dard Dueling-Bandit framework which applies to several practical scenarios such as medical treatment preferences, recommender systems, search engine optimization etc., where it is eas- ier and more effective to collect feedback for multiple options simultaneously. We develop a novel class of pairwise-subset choice model, for modelling the subset-wise winner feedback and propose three algorithms - Battling-Doubler, Battling-MultiSBM and Battling-Duel: While the first two are designed for a special class of linear-link based choice models, the third one applies to a much general class of pairwise- subset choice models with Condorcet winner. We also analyzed their regret guarantees and show the optimality of Battling-Duel proving a matching regret lower bound of $\Omega$(n log T), which (perhaps surprisingly) shows that the flexibility of playing size-k subsets does not really help to gather information faster than the corresponding dueling case (k = 2), at least for the current subsetwise feedback choice model. The efficacy of our algorithms are demonstrated through extensive experimental evaluations on a variety of synthetic and real world datasets.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Battle of Bandits %A Aadirupa Saha %A Aditya Gopalan %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-saha18a %I PMLR %P 804--813 %U https://proceedings.mlr.press/r16/saha18a.html %V R16 %X We introduce Battling-Bandits – an online learning framework where given a set of n arms, the learner needs to select a subset of k $\geq$2 arms in each round and subsequently observes a stochastic feedback indicating the winner of the round. This framework generalizes the stan- dard Dueling-Bandit framework which applies to several practical scenarios such as medical treatment preferences, recommender systems, search engine optimization etc., where it is eas- ier and more effective to collect feedback for multiple options simultaneously. We develop a novel class of pairwise-subset choice model, for modelling the subset-wise winner feedback and propose three algorithms - Battling-Doubler, Battling-MultiSBM and Battling-Duel: While the first two are designed for a special class of linear-link based choice models, the third one applies to a much general class of pairwise- subset choice models with Condorcet winner. We also analyzed their regret guarantees and show the optimality of Battling-Duel proving a matching regret lower bound of $\Omega$(n log T), which (perhaps surprisingly) shows that the flexibility of playing size-k subsets does not really help to gather information faster than the corresponding dueling case (k = 2), at least for the current subsetwise feedback choice model. The efficacy of our algorithms are demonstrated through extensive experimental evaluations on a variety of synthetic and real world datasets. %Z Reissued by PMLR on 04 October 2026.
APA
Saha, A. & Gopalan, A.. (2018). Battle of Bandits. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:804-813 Available from https://proceedings.mlr.press/r16/saha18a.html. Reissued by PMLR on 04 October 2026.

Related Material