Probabilistic inverse reinforcement learning in unknown environments

Aristide Tossou, Christos Dimitrakakis
Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence, PMLR R11:678-686, 2013.

Abstract

We consider the problem of learning by demonstration from agents acting in un- known stochastic Markov environments or games. Our aim is to estimate agent prefer- ences in order to construct improved policies for the same task that the agents are trying to solve. To do so, we extend previous prob- abilistic approaches for inverse reinforcement learning in known MDPs to the case of un- known dynamics or opponents. We do this by deriving two simplified probabilistic mod- els of the demonstrator’s policy and utility. For tractability, we use maximum a posteri- ori estimation rather than full Bayesian in- ference. Under a flat prior, this results in a convex optimisation problem. We find that the resulting algorithms are highly compet- itive against a variety of other methods for inverse reinforcement learning that do have knowledge of the dynamics.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR11-tossou13a, title = {Probabilistic inverse reinforcement learning in unknown environments}, author = {Tossou, Aristide and Dimitrakakis, Christos}, booktitle = {Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence}, pages = {678--686}, year = {2013}, editor = {Nicholson, Ann and Smyth, Padhraic}, volume = {R11}, series = {Proceedings of Machine Learning Research}, month = {12--14 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r11/main/assets/tossou13a/tossou13a.pdf}, url = {https://proceedings.mlr.press/r11/tossou13a.html}, abstract = {We consider the problem of learning by demonstration from agents acting in un- known stochastic Markov environments or games. Our aim is to estimate agent prefer- ences in order to construct improved policies for the same task that the agents are trying to solve. To do so, we extend previous prob- abilistic approaches for inverse reinforcement learning in known MDPs to the case of un- known dynamics or opponents. We do this by deriving two simplified probabilistic mod- els of the demonstrator’s policy and utility. For tractability, we use maximum a posteri- ori estimation rather than full Bayesian in- ference. Under a flat prior, this results in a convex optimisation problem. We find that the resulting algorithms are highly compet- itive against a variety of other methods for inverse reinforcement learning that do have knowledge of the dynamics.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Probabilistic inverse reinforcement learning in unknown environments %A Aristide Tossou %A Christos Dimitrakakis %B Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2013 %E Ann Nicholson %E Padhraic Smyth %F pmlr-vR11-tossou13a %I PMLR %P 678--686 %U https://proceedings.mlr.press/r11/tossou13a.html %V R11 %X We consider the problem of learning by demonstration from agents acting in un- known stochastic Markov environments or games. Our aim is to estimate agent prefer- ences in order to construct improved policies for the same task that the agents are trying to solve. To do so, we extend previous prob- abilistic approaches for inverse reinforcement learning in known MDPs to the case of un- known dynamics or opponents. We do this by deriving two simplified probabilistic mod- els of the demonstrator’s policy and utility. For tractability, we use maximum a posteri- ori estimation rather than full Bayesian in- ference. Under a flat prior, this results in a convex optimisation problem. We find that the resulting algorithms are highly compet- itive against a variety of other methods for inverse reinforcement learning that do have knowledge of the dynamics. %Z Reissued by PMLR on 04 October 2026.
APA
Tossou, A. & Dimitrakakis, C.. (2013). Probabilistic inverse reinforcement learning in unknown environments. Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R11:678-686 Available from https://proceedings.mlr.press/r11/tossou13a.html. Reissued by PMLR on 04 October 2026.

Related Material