Bayesian Optimal Control of Smoothly Parameterized Systems

Yasin Abbasi-Yadkori QUT, Csaba Szepesvari
Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, PMLR R13:810-819, 2015.

Abstract

We study Bayesian optimal control of a general class of smoothly parameterized Markov decision problems (MDPs). We propose a \textit{lazy} version of the so-called posterior sampling method, a method that goes back to Thompson and Strens, more recently studied by Osband, Russo and van Roy. While Osband et al. derived a bound on the (Bayesian) regret of this method for undiscounted total cost episodic, finite state and action problems, we consider the continuing, average cost setting with no cardinality restrictions on the state or action spaces. While in the episodic setting, it is natural to switch to a new policy at the episode-ends, in the continuing average cost framework we must introduce switching points explicitly and in a principled fashion, or the regret could grow linearly. Our lazy method introduces these switching points based on monitoring the uncertainty left about the unknown parameter. To develop a suitable and easy-to-compute uncertainty measure, we introduce a new “average local smoothness” condition, which is shown to be satisfied in common examples. Under this, and some additional mild conditions, we derive rate-optimal bounds on the regret of our algorithm. Our general approach allows us to use a single algorithm and a single analysis for a wide range of problems, such as finite MDPs or linear quadratic regulation, both being instances of smoothly parameterized MDPs. The effectiveness of our method is illustrated by means of a simulated example.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR13-qut15a, title = {{B}ayesian Optimal Control of Smoothly Parameterized Systems}, author = {QUT, Yasin Abbasi-Yadkori and Szepesvari, Csaba}, booktitle = {Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence}, pages = {810--819}, year = {2015}, editor = {Meila, Marina and Heskes, Tom}, volume = {R13}, series = {Proceedings of Machine Learning Research}, month = {12--16 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r13/main/assets/qut15a/qut15a.pdf}, url = {https://proceedings.mlr.press/r13/qut15a.html}, abstract = {We study Bayesian optimal control of a general class of smoothly parameterized Markov decision problems (MDPs). We propose a \textit{lazy} version of the so-called posterior sampling method, a method that goes back to Thompson and Strens, more recently studied by Osband, Russo and van Roy. While Osband et al. derived a bound on the (Bayesian) regret of this method for undiscounted total cost episodic, finite state and action problems, we consider the continuing, average cost setting with no cardinality restrictions on the state or action spaces. While in the episodic setting, it is natural to switch to a new policy at the episode-ends, in the continuing average cost framework we must introduce switching points explicitly and in a principled fashion, or the regret could grow linearly. Our lazy method introduces these switching points based on monitoring the uncertainty left about the unknown parameter. To develop a suitable and easy-to-compute uncertainty measure, we introduce a new “average local smoothness” condition, which is shown to be satisfied in common examples. Under this, and some additional mild conditions, we derive rate-optimal bounds on the regret of our algorithm. Our general approach allows us to use a single algorithm and a single analysis for a wide range of problems, such as finite MDPs or linear quadratic regulation, both being instances of smoothly parameterized MDPs. The effectiveness of our method is illustrated by means of a simulated example.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Bayesian Optimal Control of Smoothly Parameterized Systems %A Yasin Abbasi-Yadkori QUT %A Csaba Szepesvari %B Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2015 %E Marina Meila %E Tom Heskes %F pmlr-vR13-qut15a %I PMLR %P 810--819 %U https://proceedings.mlr.press/r13/qut15a.html %V R13 %X We study Bayesian optimal control of a general class of smoothly parameterized Markov decision problems (MDPs). We propose a \textit{lazy} version of the so-called posterior sampling method, a method that goes back to Thompson and Strens, more recently studied by Osband, Russo and van Roy. While Osband et al. derived a bound on the (Bayesian) regret of this method for undiscounted total cost episodic, finite state and action problems, we consider the continuing, average cost setting with no cardinality restrictions on the state or action spaces. While in the episodic setting, it is natural to switch to a new policy at the episode-ends, in the continuing average cost framework we must introduce switching points explicitly and in a principled fashion, or the regret could grow linearly. Our lazy method introduces these switching points based on monitoring the uncertainty left about the unknown parameter. To develop a suitable and easy-to-compute uncertainty measure, we introduce a new “average local smoothness” condition, which is shown to be satisfied in common examples. Under this, and some additional mild conditions, we derive rate-optimal bounds on the regret of our algorithm. Our general approach allows us to use a single algorithm and a single analysis for a wide range of problems, such as finite MDPs or linear quadratic regulation, both being instances of smoothly parameterized MDPs. The effectiveness of our method is illustrated by means of a simulated example. %Z Reissued by PMLR on 04 October 2026.
APA
QUT, Y.A. & Szepesvari, C.. (2015). Bayesian Optimal Control of Smoothly Parameterized Systems. Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R13:810-819 Available from https://proceedings.mlr.press/r13/qut15a.html. Reissued by PMLR on 04 October 2026.

Related Material