Pure Exploration of Multi-Armed Bandits with Heavy-Tailed Payoffs

Xiaotian Yu, Han Shao, Michael R. Lyu, Irwin King
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:936-945, 2018.

Abstract

Inspired by heavy-tailed distributions in prac- tical scenarios, we investigate the problem on pure exploration of Multi-Armed Bandits (MAB) with heavy-tailed payoffs by breaking the assumption of payoffs with sub-Gaussian noises in MAB, and assuming that stochastic payoffs from bandits are with finite p-th mo- ments, where p $\in$(1, +$\infty$). The main contri- butions in this paper are three-fold. First, we technically analyze tail probabilities of empir- ical average and truncated empirical average (TEA) for estimating expected payoffs in se- quential decisions with heavy-tailed noises via martingales. Second, we propose two effective bandit algorithms based on different prior in- formation (i.e., fixed confidence or fixed bud- get) for pure exploration of MAB generating payoffs with finite p-th moments. Third, we derive theoretical guarantees for the proposed two bandit algorithms, and demonstrate the ef- fectiveness of two algorithms in pure explo- ration of MAB with heavy-tailed payoffs in synthetic data and real-world financial data.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-yu18a, title = {Pure Exploration of Multi-Armed Bandits with Heavy-Tailed Payoffs}, author = {Yu, Xiaotian and Shao, Han and Lyu, Michael R. and King, Irwin}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {936--945}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/yu18a/yu18a.pdf}, url = {https://proceedings.mlr.press/r16/yu18a.html}, abstract = {Inspired by heavy-tailed distributions in prac- tical scenarios, we investigate the problem on pure exploration of Multi-Armed Bandits (MAB) with heavy-tailed payoffs by breaking the assumption of payoffs with sub-Gaussian noises in MAB, and assuming that stochastic payoffs from bandits are with finite p-th mo- ments, where p $\in$(1, +$\infty$). The main contri- butions in this paper are three-fold. First, we technically analyze tail probabilities of empir- ical average and truncated empirical average (TEA) for estimating expected payoffs in se- quential decisions with heavy-tailed noises via martingales. Second, we propose two effective bandit algorithms based on different prior in- formation (i.e., fixed confidence or fixed bud- get) for pure exploration of MAB generating payoffs with finite p-th moments. Third, we derive theoretical guarantees for the proposed two bandit algorithms, and demonstrate the ef- fectiveness of two algorithms in pure explo- ration of MAB with heavy-tailed payoffs in synthetic data and real-world financial data.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Pure Exploration of Multi-Armed Bandits with Heavy-Tailed Payoffs %A Xiaotian Yu %A Han Shao %A Michael R. Lyu %A Irwin King %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-yu18a %I PMLR %P 936--945 %U https://proceedings.mlr.press/r16/yu18a.html %V R16 %X Inspired by heavy-tailed distributions in prac- tical scenarios, we investigate the problem on pure exploration of Multi-Armed Bandits (MAB) with heavy-tailed payoffs by breaking the assumption of payoffs with sub-Gaussian noises in MAB, and assuming that stochastic payoffs from bandits are with finite p-th mo- ments, where p $\in$(1, +$\infty$). The main contri- butions in this paper are three-fold. First, we technically analyze tail probabilities of empir- ical average and truncated empirical average (TEA) for estimating expected payoffs in se- quential decisions with heavy-tailed noises via martingales. Second, we propose two effective bandit algorithms based on different prior in- formation (i.e., fixed confidence or fixed bud- get) for pure exploration of MAB generating payoffs with finite p-th moments. Third, we derive theoretical guarantees for the proposed two bandit algorithms, and demonstrate the ef- fectiveness of two algorithms in pure explo- ration of MAB with heavy-tailed payoffs in synthetic data and real-world financial data. %Z Reissued by PMLR on 04 October 2026.
APA
Yu, X., Shao, H., Lyu, M.R. & King, I.. (2018). Pure Exploration of Multi-Armed Bandits with Heavy-Tailed Payoffs. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:936-945 Available from https://proceedings.mlr.press/r16/yu18a.html. Reissued by PMLR on 04 October 2026.

Related Material