Dynamic Teaching in Sequential Decision Making Environments

Thomas J. Walsh, Sergiu Goschin
Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence, PMLR R10:862-871, 2012.

Abstract

We describe theoretical bounds and a practical algorithm for teaching a model by demonstration in a sequential decision making environment. Unlike previous efforts that have optimized learners that watch a teacher demonstrate a static policy, we focus on the teacher as a decision maker who can dynamically choose different policies to teach different parts of the environment. We develop several teaching frameworks based on previously defined supervised protocols, such as Teaching Dimension, extending them to handle noise and sequences of inputs encountered in an MDP.We provide theoretical bounds on the learnability of several important model classes in this setting and suggest a practical algorithm for dynamic teaching.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR10-walsh12a, title = {Dynamic Teaching in Sequential Decision Making Environments}, author = {Walsh, Thomas J. and Goschin, Sergiu}, booktitle = {Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence}, pages = {862--871}, year = {2012}, editor = {de Freitas, Nando and Murphy, Kevin}, volume = {R10}, series = {Proceedings of Machine Learning Research}, month = {14--18 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r10/main/assets/walsh12a/walsh12a.pdf}, url = {https://proceedings.mlr.press/r10/walsh12a.html}, abstract = {We describe theoretical bounds and a practical algorithm for teaching a model by demonstration in a sequential decision making environment. Unlike previous efforts that have optimized learners that watch a teacher demonstrate a static policy, we focus on the teacher as a decision maker who can dynamically choose different policies to teach different parts of the environment. We develop several teaching frameworks based on previously defined supervised protocols, such as Teaching Dimension, extending them to handle noise and sequences of inputs encountered in an MDP.We provide theoretical bounds on the learnability of several important model classes in this setting and suggest a practical algorithm for dynamic teaching.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Dynamic Teaching in Sequential Decision Making Environments %A Thomas J. Walsh %A Sergiu Goschin %B Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2012 %E Nando de Freitas %E Kevin Murphy %F pmlr-vR10-walsh12a %I PMLR %P 862--871 %U https://proceedings.mlr.press/r10/walsh12a.html %V R10 %X We describe theoretical bounds and a practical algorithm for teaching a model by demonstration in a sequential decision making environment. Unlike previous efforts that have optimized learners that watch a teacher demonstrate a static policy, we focus on the teacher as a decision maker who can dynamically choose different policies to teach different parts of the environment. We develop several teaching frameworks based on previously defined supervised protocols, such as Teaching Dimension, extending them to handle noise and sequences of inputs encountered in an MDP.We provide theoretical bounds on the learnability of several important model classes in this setting and suggest a practical algorithm for dynamic teaching. %Z Reissued by PMLR on 04 October 2026.
APA
Walsh, T.J. & Goschin, S.. (2012). Dynamic Teaching in Sequential Decision Making Environments. Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R10:862-871 Available from https://proceedings.mlr.press/r10/walsh12a.html. Reissued by PMLR on 04 October 2026.

Related Material