Cascading Bandits for Large-Scale Recommendation Problems

Shi Zong, Hao Ni, Kenny Sung, Rosemary Ke, Zheng Wen Adobe Research, Branislav Kveton Adobe Research
Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, PMLR R14:286-295, 2016.

Abstract

Most recommender systems recommend a list of items. The user examines the list, from the first item to the last, and often chooses the first attractive item and does not examine the rest. This type of user behavior can be modeled by the cascade model. In this work, we study cascading bandits, an online learning variant of the cascade model where the goal is to recommend K most attractive items from a large set of L candidate items. We propose two algorithms for solving this problem, which are based on the idea of linear generalization. The key idea in our solutions is that we learn a predictor of the attraction probabilities of items from their features, as opposing to learning the attraction probability of each item independently as in the existing work. This results in practical learning algorithms whose regret does not depend on the number of items L. We bound the regret of one algorithm and comprehensively evaluate the other on a range of recommendation problems. The algorithm performs well and outperforms all baselines.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR14-zong16a, title = {Cascading Bandits for Large-Scale Recommendation Problems}, author = {Zong, Shi and Ni, Hao and Sung, Kenny and Ke, Rosemary and Research, Zheng Wen Adobe and Research, Branislav Kveton Adobe}, booktitle = {Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence}, pages = {286--295}, year = {2016}, editor = {Ihler, Alexander and Janzing, Dominik}, volume = {R14}, series = {Proceedings of Machine Learning Research}, month = {25--29 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r14/main/assets/zong16a/zong16a.pdf}, url = {https://proceedings.mlr.press/r14/zong16a.html}, abstract = {Most recommender systems recommend a list of items. The user examines the list, from the first item to the last, and often chooses the first attractive item and does not examine the rest. This type of user behavior can be modeled by the cascade model. In this work, we study cascading bandits, an online learning variant of the cascade model where the goal is to recommend K most attractive items from a large set of L candidate items. We propose two algorithms for solving this problem, which are based on the idea of linear generalization. The key idea in our solutions is that we learn a predictor of the attraction probabilities of items from their features, as opposing to learning the attraction probability of each item independently as in the existing work. This results in practical learning algorithms whose regret does not depend on the number of items L. We bound the regret of one algorithm and comprehensively evaluate the other on a range of recommendation problems. The algorithm performs well and outperforms all baselines.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Cascading Bandits for Large-Scale Recommendation Problems %A Shi Zong %A Hao Ni %A Kenny Sung %A Rosemary Ke %A Zheng Wen Adobe Research %A Branislav Kveton Adobe Research %B Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2016 %E Alexander Ihler %E Dominik Janzing %F pmlr-vR14-zong16a %I PMLR %P 286--295 %U https://proceedings.mlr.press/r14/zong16a.html %V R14 %X Most recommender systems recommend a list of items. The user examines the list, from the first item to the last, and often chooses the first attractive item and does not examine the rest. This type of user behavior can be modeled by the cascade model. In this work, we study cascading bandits, an online learning variant of the cascade model where the goal is to recommend K most attractive items from a large set of L candidate items. We propose two algorithms for solving this problem, which are based on the idea of linear generalization. The key idea in our solutions is that we learn a predictor of the attraction probabilities of items from their features, as opposing to learning the attraction probability of each item independently as in the existing work. This results in practical learning algorithms whose regret does not depend on the number of items L. We bound the regret of one algorithm and comprehensively evaluate the other on a range of recommendation problems. The algorithm performs well and outperforms all baselines. %Z Reissued by PMLR on 04 October 2026.
APA
Zong, S., Ni, H., Sung, K., Ke, R., Research, Z.W.A. & Research, B.K.A.. (2016). Cascading Bandits for Large-Scale Recommendation Problems. Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R14:286-295 Available from https://proceedings.mlr.press/r14/zong16a.html. Reissued by PMLR on 04 October 2026.

Related Material