[edit]
Fast Policy Learning through Imitation and Reinforcement
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:844-854, 2018.
Abstract
Imitation learning (IL) consists of a set of tools that leverage expert demonstrations to quickly learn policies. However, if the expert is subop- timal, IL can yield policies with inferior per- formance compared to reinforcement learning (RL). In this paper, we aim to provide an algo- rithm that combines the best aspects of RL and IL. We accomplish this by formulating sev- eral popular RL and IL algorithms in a com- mon mirror descent framework, showing that these algorithms can be viewed as a variation on a single approach. We then propose LOKI, a strategy for policy learning that first performs a small but random number of IL iterations be- fore switching to a policy gradient RL method. We show that if the switching time is prop- erly randomized, LOKI can learn to outperform a suboptimal expert and converge faster than running policy gradient from scratch. Finally, we evaluate the performance of LOKI experi- mentally in several simulated environments.