[edit]
Inverse Reinforcement Learning via Deep Gaussian Process
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:471-480, 2017.
Abstract
We propose a new approach to inverse rein- forcement learning (IRL) based on the deep Gaussian process (deep GP) model, which is capable of learning complicated reward struc- tures with few demonstrations. Our model stacks multiple latent GP layers to learn ab- stract representations of the state feature space, which is linked to the demonstrations through the Maximum Entropy learning framework. In- corporating the IRL engine into the nonlinear latent structure renders existing deep GP infer- ence approaches intractable. To tackle this, we develop a non-standard variational approxima- tion framework which extends previous infer- ence schemes. This allows for approximate Bayesian treatment of the feature space and guards against overfitting. Carrying out rep- resentation and inverse reinforcement learn- ing simultaneously within our model outper- forms state-of-the-art approaches, as we demon- strate with experiments on standard bench- marks (“object world”,“highway driving”) and a new benchmark (“binary world”).