Bandits with Side Observations: Bounded vs. Logarithmic Regret

Rémy Degenne, Evrard Garcelon, Vianney Perchet
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:466-475, 2018.

Abstract

We consider the classical stochastic multi- armed bandit but where, from time to time and roughly with frequency $\epsilon$, an extra observation is gathered by the agent for free. We prove that, no matter how small $\epsilon$ is the agent can ensure a regret uniformly bounded in time. More precisely, we construct an algorithm with a regret smaller than P i log(1/$\epsilon$) $\Delta$i , up to multi- plicative constant and log log terms. We also prove a matching lower-bound, stating that no reasonable algorithm can outperform this quantity.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-degenne18a, title = {Bandits with Side Observations: Bounded vs. Logarithmic Regret}, author = {Degenne, R{\'e}my and Garcelon, Evrard and Perchet, Vianney}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {466--475}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/degenne18a/degenne18a.pdf}, url = {https://proceedings.mlr.press/r16/degenne18a.html}, abstract = {We consider the classical stochastic multi- armed bandit but where, from time to time and roughly with frequency $\epsilon$, an extra observation is gathered by the agent for free. We prove that, no matter how small $\epsilon$ is the agent can ensure a regret uniformly bounded in time. More precisely, we construct an algorithm with a regret smaller than P i log(1/$\epsilon$) $\Delta$i , up to multi- plicative constant and log log terms. We also prove a matching lower-bound, stating that no reasonable algorithm can outperform this quantity.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Bandits with Side Observations: Bounded vs. Logarithmic Regret %A Rémy Degenne %A Evrard Garcelon %A Vianney Perchet %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-degenne18a %I PMLR %P 466--475 %U https://proceedings.mlr.press/r16/degenne18a.html %V R16 %X We consider the classical stochastic multi- armed bandit but where, from time to time and roughly with frequency $\epsilon$, an extra observation is gathered by the agent for free. We prove that, no matter how small $\epsilon$ is the agent can ensure a regret uniformly bounded in time. More precisely, we construct an algorithm with a regret smaller than P i log(1/$\epsilon$) $\Delta$i , up to multi- plicative constant and log log terms. We also prove a matching lower-bound, stating that no reasonable algorithm can outperform this quantity. %Z Reissued by PMLR on 04 October 2026.
APA
Degenne, R., Garcelon, E. & Perchet, V.. (2018). Bandits with Side Observations: Bounded vs. Logarithmic Regret. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:466-475 Available from https://proceedings.mlr.press/r16/degenne18a.html. Reissued by PMLR on 04 October 2026.

Related Material