Thompson Sampling is Asymptotically Optimal in General Environments

Jan Leike, Tor Lattimore, Laurent Orseau, Marcus Hutter
Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, PMLR R14:88-97, 2016.

Abstract

We discuss a variant of Thompson sampling for nonparametric reinforcement learning in a countable classes of general stochastic environments. These environments can be non-Markov, non-ergodic, and partially observable. We show that Thompson sampling learns the environment classin the sense that (1) asymptotically its value converges to the optimal value in mean and (2) given a recoverability assumption regret is sublinear.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR14-leike16a, title = {{T}hompson Sampling is Asymptotically Optimal in General Environments}, author = {Leike, Jan and Lattimore, Tor and Orseau, Laurent and Hutter, Marcus}, booktitle = {Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence}, pages = {88--97}, year = {2016}, editor = {Ihler, Alexander and Janzing, Dominik}, volume = {R14}, series = {Proceedings of Machine Learning Research}, month = {25--29 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r14/main/assets/leike16a/leike16a.pdf}, url = {https://proceedings.mlr.press/r14/leike16a.html}, abstract = {We discuss a variant of Thompson sampling for nonparametric reinforcement learning in a countable classes of general stochastic environments. These environments can be non-Markov, non-ergodic, and partially observable. We show that Thompson sampling learns the environment classin the sense that (1) asymptotically its value converges to the optimal value in mean and (2) given a recoverability assumption regret is sublinear.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Thompson Sampling is Asymptotically Optimal in General Environments %A Jan Leike %A Tor Lattimore %A Laurent Orseau %A Marcus Hutter %B Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2016 %E Alexander Ihler %E Dominik Janzing %F pmlr-vR14-leike16a %I PMLR %P 88--97 %U https://proceedings.mlr.press/r14/leike16a.html %V R14 %X We discuss a variant of Thompson sampling for nonparametric reinforcement learning in a countable classes of general stochastic environments. These environments can be non-Markov, non-ergodic, and partially observable. We show that Thompson sampling learns the environment classin the sense that (1) asymptotically its value converges to the optimal value in mean and (2) given a recoverability assumption regret is sublinear. %Z Reissued by PMLR on 04 October 2026.
APA
Leike, J., Lattimore, T., Orseau, L. & Hutter, M.. (2016). Thompson Sampling is Asymptotically Optimal in General Environments. Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R14:88-97 Available from https://proceedings.mlr.press/r14/leike16a.html. Reissued by PMLR on 04 October 2026.

Related Material