Fast Policy Learning through Imitation and Reinforcement

Ching-An Cheng, Xinyan Yan, Nolan Wagener, Byron Boots
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:844-854, 2018.

Abstract

Imitation learning (IL) consists of a set of tools that leverage expert demonstrations to quickly learn policies. However, if the expert is subop- timal, IL can yield policies with inferior per- formance compared to reinforcement learning (RL). In this paper, we aim to provide an algo- rithm that combines the best aspects of RL and IL. We accomplish this by formulating sev- eral popular RL and IL algorithms in a com- mon mirror descent framework, showing that these algorithms can be viewed as a variation on a single approach. We then propose LOKI, a strategy for policy learning that first performs a small but random number of IL iterations be- fore switching to a policy gradient RL method. We show that if the switching time is prop- erly randomized, LOKI can learn to outperform a suboptimal expert and converge faster than running policy gradient from scratch. Finally, we evaluate the performance of LOKI experi- mentally in several simulated environments.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-cheng18a, title = {Fast Policy Learning through Imitation and Reinforcement}, author = {Cheng, Ching-An and Yan, Xinyan and Wagener, Nolan and Boots, Byron}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {844--854}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/cheng18a/cheng18a.pdf}, url = {https://proceedings.mlr.press/r16/cheng18a.html}, abstract = {Imitation learning (IL) consists of a set of tools that leverage expert demonstrations to quickly learn policies. However, if the expert is subop- timal, IL can yield policies with inferior per- formance compared to reinforcement learning (RL). In this paper, we aim to provide an algo- rithm that combines the best aspects of RL and IL. We accomplish this by formulating sev- eral popular RL and IL algorithms in a com- mon mirror descent framework, showing that these algorithms can be viewed as a variation on a single approach. We then propose LOKI, a strategy for policy learning that first performs a small but random number of IL iterations be- fore switching to a policy gradient RL method. We show that if the switching time is prop- erly randomized, LOKI can learn to outperform a suboptimal expert and converge faster than running policy gradient from scratch. Finally, we evaluate the performance of LOKI experi- mentally in several simulated environments.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Fast Policy Learning through Imitation and Reinforcement %A Ching-An Cheng %A Xinyan Yan %A Nolan Wagener %A Byron Boots %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-cheng18a %I PMLR %P 844--854 %U https://proceedings.mlr.press/r16/cheng18a.html %V R16 %X Imitation learning (IL) consists of a set of tools that leverage expert demonstrations to quickly learn policies. However, if the expert is subop- timal, IL can yield policies with inferior per- formance compared to reinforcement learning (RL). In this paper, we aim to provide an algo- rithm that combines the best aspects of RL and IL. We accomplish this by formulating sev- eral popular RL and IL algorithms in a com- mon mirror descent framework, showing that these algorithms can be viewed as a variation on a single approach. We then propose LOKI, a strategy for policy learning that first performs a small but random number of IL iterations be- fore switching to a policy gradient RL method. We show that if the switching time is prop- erly randomized, LOKI can learn to outperform a suboptimal expert and converge faster than running policy gradient from scratch. Finally, we evaluate the performance of LOKI experi- mentally in several simulated environments. %Z Reissued by PMLR on 04 October 2026.
APA
Cheng, C., Yan, X., Wagener, N. & Boots, B.. (2018). Fast Policy Learning through Imitation and Reinforcement. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:844-854 Available from https://proceedings.mlr.press/r16/cheng18a.html. Reissued by PMLR on 04 October 2026.

Related Material