Model Regularization for Stable Sample Rollouts

Erik Talvitie Franklin & Marshall College
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:568-577, 2014.

Abstract

When an imperfect model is used to generate sample rollouts, its errors tend to compound – a flawed sample is given as input to the model, which causes more errors, and so on. This presents a barrier to applying rollout-based plan- ning algorithms to learned models. To ad- dress this issue, a training methodology called “hallucinated replay” is introduced, which adds samples from the model into the training data, thereby training the model to produce sensible predictions when its own samples are given as input. Capabilities and limitations of this ap- proach are studied empirically. In several exam- ples hallucinated replay allows effective planning with imperfect models while models trained us- ing only real experience fail dramatically.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-college14a, title = {Model Regularization for Stable Sample Rollouts}, author = {College, Erik Talvitie Franklin \& Marshall}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {568--577}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/college14a/college14a.pdf}, url = {https://proceedings.mlr.press/r12/college14a.html}, abstract = {When an imperfect model is used to generate sample rollouts, its errors tend to compound – a flawed sample is given as input to the model, which causes more errors, and so on. This presents a barrier to applying rollout-based plan- ning algorithms to learned models. To ad- dress this issue, a training methodology called “hallucinated replay” is introduced, which adds samples from the model into the training data, thereby training the model to produce sensible predictions when its own samples are given as input. Capabilities and limitations of this ap- proach are studied empirically. In several exam- ples hallucinated replay allows effective planning with imperfect models while models trained us- ing only real experience fail dramatically.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Model Regularization for Stable Sample Rollouts %A Erik Talvitie Franklin & Marshall College %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-college14a %I PMLR %P 568--577 %U https://proceedings.mlr.press/r12/college14a.html %V R12 %X When an imperfect model is used to generate sample rollouts, its errors tend to compound – a flawed sample is given as input to the model, which causes more errors, and so on. This presents a barrier to applying rollout-based plan- ning algorithms to learned models. To ad- dress this issue, a training methodology called “hallucinated replay” is introduced, which adds samples from the model into the training data, thereby training the model to produce sensible predictions when its own samples are given as input. Capabilities and limitations of this ap- proach are studied empirically. In several exam- ples hallucinated replay allows effective planning with imperfect models while models trained us- ing only real experience fail dramatically. %Z Reissued by PMLR on 04 October 2026.
APA
College, E.T.F.&.M.. (2014). Model Regularization for Stable Sample Rollouts. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:568-577 Available from https://proceedings.mlr.press/r12/college14a.html. Reissued by PMLR on 04 October 2026.

Related Material