Variational Imitation Learning with Diverse-quality Demonstrations

Voot Tangkaratt; Bo Han; Mohammad Emtiyaz Khan; Masashi Sugiyama

Variational Imitation Learning with Diverse-quality Demonstrations

Voot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi Sugiyama

Proceedings of the 37th International Conference on Machine Learning, PMLR 119:9407-9417, 2020.

Abstract

Learning from demonstrations can be challenging when the quality of demonstrations is diverse, and even more so when the quality is unknown and there is no additional information to estimate the quality. We propose a new method for imitation learning in such scenarios. We show that simple quality-estimation approaches might fail due to compounding error, and fix this issue by jointly estimating both the quality and reward using a variational approach. Our method is easy to implement within reinforcement-learning frameworks and also achieves state-of-the-art performance on continuous-control benchmarks.Our work enables scalable and data-efficient imitation learning under more realistic settings than before.

Cite this Paper

BibTeX


@InProceedings{pmlr-v119-tangkaratt20a,
  title = 	 {Variational Imitation Learning with Diverse-quality Demonstrations},
  author =       {Tangkaratt, Voot and Han, Bo and Khan, Mohammad Emtiyaz and Sugiyama, Masashi},
  booktitle = 	 {Proceedings of the 37th International Conference on Machine Learning},
  pages = 	 {9407--9417},
  year = 	 {2020},
  editor = 	 {III, Hal Daumé and Singh, Aarti},
  volume = 	 {119},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {13--18 Jul},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v119/tangkaratt20a/tangkaratt20a.pdf},
  url = 	 {https://proceedings.mlr.press/v119/tangkaratt20a.html},
  abstract = 	 {Learning from demonstrations can be challenging when the quality of demonstrations is diverse, and even more so when the quality is unknown and there is no additional information to estimate the quality. We propose a new method for imitation learning in such scenarios. We show that simple quality-estimation approaches might fail due to compounding error, and fix this issue by jointly estimating both the quality and reward using a variational approach. Our method is easy to implement within reinforcement-learning frameworks and also achieves state-of-the-art performance on continuous-control benchmarks.Our work enables scalable and data-efficient imitation learning under more realistic settings than before.}
}

Endnote

%0 Conference Paper
%T Variational Imitation Learning with Diverse-quality Demonstrations
%A Voot Tangkaratt
%A Bo Han
%A Mohammad Emtiyaz Khan
%A Masashi Sugiyama
%B Proceedings of the 37th International Conference on Machine Learning
%C Proceedings of Machine Learning Research
%D 2020
%E Hal Daumé III
%E Aarti Singh	
%F pmlr-v119-tangkaratt20a
%I PMLR
%P 9407--9417
%U https://proceedings.mlr.press/v119/tangkaratt20a.html
%V 119
%X Learning from demonstrations can be challenging when the quality of demonstrations is diverse, and even more so when the quality is unknown and there is no additional information to estimate the quality. We propose a new method for imitation learning in such scenarios. We show that simple quality-estimation approaches might fail due to compounding error, and fix this issue by jointly estimating both the quality and reward using a variational approach. Our method is easy to implement within reinforcement-learning frameworks and also achieves state-of-the-art performance on continuous-control benchmarks.Our work enables scalable and data-efficient imitation learning under more realistic settings than before.

APA


Tangkaratt, V., Han, B., Khan, M.E. & Sugiyama, M.. (2020). Variational Imitation Learning with Diverse-quality Demonstrations. Proceedings of the 37th International Conference on Machine Learning, in Proceedings of Machine Learning Research 119:9407-9417 Available from https://proceedings.mlr.press/v119/tangkaratt20a.html.

Variational Imitation Learning with Diverse-quality Demonstrations

Abstract

Cite this Paper

Related Material