[edit]
Adversarial Inverse Optimal Control for General Imitation Learning Losses and Embodiment Transfer
Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, PMLR R14:326-335, 2016.
Abstract
We develop a general framework for inverse optimal control that distinguishes between rationalizing demonstrated behavior and imitating inductively inferred behavior. This enables learning for more general imitative evaluation measures and differences between the capabilities of the demonstrator and those of the learner (i.e., differences in embodiment). Our formulation takes the form of a zero-sum game between a predictor attempting to minimize an imitative loss measure, and an adversary attempting to maximize the loss by approximating the demonstrated examples in limited ways. We establish the consistency and generalization guarantees of this approach and il- lustrate its benefits on real and synthetic imitation learning tasks.