Per-decision Multi-step Temporal Difference Learning with Control Variates

Kristopher De Asis, Richard S. Sutton
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:785-793, 2018.

Abstract

Multi-step temporal difference (TD) learning is an important approach in reinforcement learning, as it unifies one-step TD learning with Monte Carlo methods in a way where intermediate algorithms can outperform ei- ther extreme. They address a bias-variance trade off between reliance on current estimates, which could be poor, and incorporating longer sampled reward sequences into the updates. Especially in the off-policy setting, where the agent aims to learn about a policy different from the one generating its behaviour, the vari- ance in the updates can cause learning to di- verge as the number of sampled rewards used in the estimates increases. In this paper, we in- troduce per-decision control variates for multi- step TD algorithms, and compare them to ex- isting methods. Our results show that includ- ing the control variates can greatly improve performance on both on and off-policy multi- step temporal difference learning tasks.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-de-asis18a, title = {Per-decision Multi-step Temporal Difference Learning with Control Variates}, author = {De Asis, Kristopher and Sutton, Richard S.}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {785--793}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/de-asis18a/de-asis18a.pdf}, url = {https://proceedings.mlr.press/r16/de-asis18a.html}, abstract = {Multi-step temporal difference (TD) learning is an important approach in reinforcement learning, as it unifies one-step TD learning with Monte Carlo methods in a way where intermediate algorithms can outperform ei- ther extreme. They address a bias-variance trade off between reliance on current estimates, which could be poor, and incorporating longer sampled reward sequences into the updates. Especially in the off-policy setting, where the agent aims to learn about a policy different from the one generating its behaviour, the vari- ance in the updates can cause learning to di- verge as the number of sampled rewards used in the estimates increases. In this paper, we in- troduce per-decision control variates for multi- step TD algorithms, and compare them to ex- isting methods. Our results show that includ- ing the control variates can greatly improve performance on both on and off-policy multi- step temporal difference learning tasks.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Per-decision Multi-step Temporal Difference Learning with Control Variates %A Kristopher De Asis %A Richard S. Sutton %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-de-asis18a %I PMLR %P 785--793 %U https://proceedings.mlr.press/r16/de-asis18a.html %V R16 %X Multi-step temporal difference (TD) learning is an important approach in reinforcement learning, as it unifies one-step TD learning with Monte Carlo methods in a way where intermediate algorithms can outperform ei- ther extreme. They address a bias-variance trade off between reliance on current estimates, which could be poor, and incorporating longer sampled reward sequences into the updates. Especially in the off-policy setting, where the agent aims to learn about a policy different from the one generating its behaviour, the vari- ance in the updates can cause learning to di- verge as the number of sampled rewards used in the estimates increases. In this paper, we in- troduce per-decision control variates for multi- step TD algorithms, and compare them to ex- isting methods. Our results show that includ- ing the control variates can greatly improve performance on both on and off-policy multi- step temporal difference learning tasks. %Z Reissued by PMLR on 04 October 2026.
APA
De Asis, K. & Sutton, R.S.. (2018). Per-decision Multi-step Temporal Difference Learning with Control Variates. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:785-793 Available from https://proceedings.mlr.press/r16/de-asis18a.html. Reissued by PMLR on 04 October 2026.

Related Material