[edit]
Per-decision Multi-step Temporal Difference Learning with Control Variates
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:785-793, 2018.
Abstract
Multi-step temporal difference (TD) learning is an important approach in reinforcement learning, as it unifies one-step TD learning with Monte Carlo methods in a way where intermediate algorithms can outperform ei- ther extreme. They address a bias-variance trade off between reliance on current estimates, which could be poor, and incorporating longer sampled reward sequences into the updates. Especially in the off-policy setting, where the agent aims to learn about a policy different from the one generating its behaviour, the vari- ance in the updates can cause learning to di- verge as the number of sampled rewards used in the estimates increases. In this paper, we in- troduce per-decision control variates for multi- step TD algorithms, and compare them to ex- isting methods. Our results show that includ- ing the control variates can greatly improve performance on both on and off-policy multi- step temporal difference learning tasks.