[edit]
Model Regularization for Stable Sample Rollouts
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:568-577, 2014.
Abstract
When an imperfect model is used to generate sample rollouts, its errors tend to compound – a flawed sample is given as input to the model, which causes more errors, and so on. This presents a barrier to applying rollout-based plan- ning algorithms to learned models. To ad- dress this issue, a training methodology called “hallucinated replay” is introduced, which adds samples from the model into the training data, thereby training the model to produce sensible predictions when its own samples are given as input. Capabilities and limitations of this ap- proach are studied empirically. In several exam- ples hallucinated replay allows effective planning with imperfect models while models trained us- ing only real experience fail dramatically.