Learning and Planning with Timing Information in Markov Decision Processes

Pierre-Luc Bacon McGill University, Borja Balle McGill University, Doina Precup McGill University
Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, PMLR R13:820-829, 2015.

Abstract

We consider the problem of learning and planning in Markov decision processes with temporally extended actions represented in the options framework. We propose to use predictions about the duration of extended actions to represent the state and show that this leads to a compact predictive state representation model independent of the set of primitive actions. Then we develop a consistent and efficient spectral learning algorithm for such models. Using just the timing information to represent states allows for faster improvement in the planning performance. We illustrate our approach with experiments in both synthetic and robot navigation domains.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR13-university15s, title = {Learning and Planning with Timing Information in {M}arkov Decision Processes}, author = {University, Pierre-Luc Bacon McGill and University, Borja Balle McGill and University, Doina Precup McGill}, booktitle = {Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence}, pages = {820--829}, year = {2015}, editor = {Meila, Marina and Heskes, Tom}, volume = {R13}, series = {Proceedings of Machine Learning Research}, month = {12--16 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r13/main/assets/university15s/university15s.pdf}, url = {https://proceedings.mlr.press/r13/university15s.html}, abstract = {We consider the problem of learning and planning in Markov decision processes with temporally extended actions represented in the options framework. We propose to use predictions about the duration of extended actions to represent the state and show that this leads to a compact predictive state representation model independent of the set of primitive actions. Then we develop a consistent and efficient spectral learning algorithm for such models. Using just the timing information to represent states allows for faster improvement in the planning performance. We illustrate our approach with experiments in both synthetic and robot navigation domains.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Learning and Planning with Timing Information in Markov Decision Processes %A Pierre-Luc Bacon McGill University %A Borja Balle McGill University %A Doina Precup McGill University %B Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2015 %E Marina Meila %E Tom Heskes %F pmlr-vR13-university15s %I PMLR %P 820--829 %U https://proceedings.mlr.press/r13/university15s.html %V R13 %X We consider the problem of learning and planning in Markov decision processes with temporally extended actions represented in the options framework. We propose to use predictions about the duration of extended actions to represent the state and show that this leads to a compact predictive state representation model independent of the set of primitive actions. Then we develop a consistent and efficient spectral learning algorithm for such models. Using just the timing information to represent states allows for faster improvement in the planning performance. We illustrate our approach with experiments in both synthetic and robot navigation domains. %Z Reissued by PMLR on 04 October 2026.
APA
University, P.B.M., University, B.B.M. & University, D.P.M.. (2015). Learning and Planning with Timing Information in Markov Decision Processes. Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R13:820-829 Available from https://proceedings.mlr.press/r13/university15s.html. Reissued by PMLR on 04 October 2026.

Related Material