A Bayesian Sampling Approach to Exploration in Reinforcement Learning

Michael Littman, Lihong Li, Ali Nouri, David Wingate, John Asmuth
Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, PMLR R7:339-346, 2009.

Abstract

We present a modular approach to reinforcement learning that uses a Bayesian representation of the uncertainty over models. The approach, BOSS (Best of Sampled Set), drives exploration by sampling multiple models from the posterior and selecting actions optimistically. It extends previous work by providing a rule for deciding when to resample and how to combine the models. We show that our algorithm achieves nearoptimal reward with high probability with a sample complexity that is low relative to the speed at which the posterior distribution converges during learning. We demonstrate that BOSS performs quite favorably compared to state-of-the-art reinforcement-learning approaches and illustrate its flexibility by pairing it with a non-parametric model that generalizes across states.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR7-littman09a, title = {A {B}ayesian Sampling Approach to Exploration in Reinforcement Learning}, author = {Littman, Michael and Li, Lihong and Nouri, Ali and Wingate, David and Asmuth, John}, booktitle = {Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence}, pages = {339--346}, year = {2009}, editor = {Bilmes, Jeff and Ng, Andrew Y.}, volume = {R7}, series = {Proceedings of Machine Learning Research}, month = {18--21 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r7/main/assets/littman09a/littman09a.pdf}, url = {https://proceedings.mlr.press/r7/littman09a.html}, abstract = {We present a modular approach to reinforcement learning that uses a Bayesian representation of the uncertainty over models. The approach, BOSS (Best of Sampled Set), drives exploration by sampling multiple models from the posterior and selecting actions optimistically. It extends previous work by providing a rule for deciding when to resample and how to combine the models. We show that our algorithm achieves nearoptimal reward with high probability with a sample complexity that is low relative to the speed at which the posterior distribution converges during learning. We demonstrate that BOSS performs quite favorably compared to state-of-the-art reinforcement-learning approaches and illustrate its flexibility by pairing it with a non-parametric model that generalizes across states.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T A Bayesian Sampling Approach to Exploration in Reinforcement Learning %A Michael Littman %A Lihong Li %A Ali Nouri %A David Wingate %A John Asmuth %B Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2009 %E Jeff Bilmes %E Andrew Y. Ng %F pmlr-vR7-littman09a %I PMLR %P 339--346 %U https://proceedings.mlr.press/r7/littman09a.html %V R7 %X We present a modular approach to reinforcement learning that uses a Bayesian representation of the uncertainty over models. The approach, BOSS (Best of Sampled Set), drives exploration by sampling multiple models from the posterior and selecting actions optimistically. It extends previous work by providing a rule for deciding when to resample and how to combine the models. We show that our algorithm achieves nearoptimal reward with high probability with a sample complexity that is low relative to the speed at which the posterior distribution converges during learning. We demonstrate that BOSS performs quite favorably compared to state-of-the-art reinforcement-learning approaches and illustrate its flexibility by pairing it with a non-parametric model that generalizes across states. %Z Reissued by PMLR on 04 October 2026.
APA
Littman, M., Li, L., Nouri, A., Wingate, D. & Asmuth, J.. (2009). A Bayesian Sampling Approach to Exploration in Reinforcement Learning. Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R7:339-346 Available from https://proceedings.mlr.press/r7/littman09a.html. Reissued by PMLR on 04 October 2026.

Related Material