Comparing Direct and Indirect Temporal-Difference Methods for Estimating the Variance of the Return

Craig Sherstan, Dylan R. Ashley, Brendan Bennett, Kenny Young, Adam White, Martha White, Richard S. Sutton
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:62-71, 2018.

Abstract

Temporal-difference (TD) learning methods are widely used in reinforcement learning to estimate the expected return for each state, without a model, because of their significant advantages in computational and data effi- ciency. For many applications involving risk mitigation, it would also be useful to estimate the variance of the return by TD methods. In this paper, we describe a way of doing this that is substantially simpler than those proposed by Tamar, Di Castro, and Mannor in 2012, or those proposed by White and White in 2016. We show that two TD learners operating in series can learn expectation and variance esti- mates. The trick is to use the square of the TD error of the expectation learner as the reward of the variance learner, and the square of the ex- pectation learner’s discount rate as the discount rate of the variance learner. With these two modifications, the variance learning problem becomes a conventional TD learning problem to which standard theoretical results can be ap- plied. Our formal results are limited to the ta- ble lookup case, for which our method is still novel, but the extension to function approxi- mation is immediate, and we provide some em- pirical results for the linear function approx- imation case. Our experimental results show that our direct method behaves just as well as a comparable indirect method, but is generally more robust.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR16-sherstan18a, title = {Comparing Direct and Indirect Temporal-Difference Methods for Estimating the Variance of the Return}, author = {Sherstan, Craig and Ashley, Dylan R. and Bennett, Brendan and Young, Kenny and White, Adam and White, Martha and Sutton, Richard S.}, booktitle = {Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence}, pages = {62--71}, year = {2018}, editor = {Globerson, Amir and Silva, Ricardo}, volume = {R16}, series = {Proceedings of Machine Learning Research}, month = {06--10 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r16/main/assets/sherstan18a/sherstan18a.pdf}, url = {https://proceedings.mlr.press/r16/sherstan18a.html}, abstract = {Temporal-difference (TD) learning methods are widely used in reinforcement learning to estimate the expected return for each state, without a model, because of their significant advantages in computational and data effi- ciency. For many applications involving risk mitigation, it would also be useful to estimate the variance of the return by TD methods. In this paper, we describe a way of doing this that is substantially simpler than those proposed by Tamar, Di Castro, and Mannor in 2012, or those proposed by White and White in 2016. We show that two TD learners operating in series can learn expectation and variance esti- mates. The trick is to use the square of the TD error of the expectation learner as the reward of the variance learner, and the square of the ex- pectation learner’s discount rate as the discount rate of the variance learner. With these two modifications, the variance learning problem becomes a conventional TD learning problem to which standard theoretical results can be ap- plied. Our formal results are limited to the ta- ble lookup case, for which our method is still novel, but the extension to function approxi- mation is immediate, and we provide some em- pirical results for the linear function approx- imation case. Our experimental results show that our direct method behaves just as well as a comparable indirect method, but is generally more robust.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Comparing Direct and Indirect Temporal-Difference Methods for Estimating the Variance of the Return %A Craig Sherstan %A Dylan R. Ashley %A Brendan Bennett %A Kenny Young %A Adam White %A Martha White %A Richard S. Sutton %B Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2018 %E Amir Globerson %E Ricardo Silva %F pmlr-vR16-sherstan18a %I PMLR %P 62--71 %U https://proceedings.mlr.press/r16/sherstan18a.html %V R16 %X Temporal-difference (TD) learning methods are widely used in reinforcement learning to estimate the expected return for each state, without a model, because of their significant advantages in computational and data effi- ciency. For many applications involving risk mitigation, it would also be useful to estimate the variance of the return by TD methods. In this paper, we describe a way of doing this that is substantially simpler than those proposed by Tamar, Di Castro, and Mannor in 2012, or those proposed by White and White in 2016. We show that two TD learners operating in series can learn expectation and variance esti- mates. The trick is to use the square of the TD error of the expectation learner as the reward of the variance learner, and the square of the ex- pectation learner’s discount rate as the discount rate of the variance learner. With these two modifications, the variance learning problem becomes a conventional TD learning problem to which standard theoretical results can be ap- plied. Our formal results are limited to the ta- ble lookup case, for which our method is still novel, but the extension to function approxi- mation is immediate, and we provide some em- pirical results for the linear function approx- imation case. Our experimental results show that our direct method behaves just as well as a comparable indirect method, but is generally more robust. %Z Reissued by PMLR on 04 October 2026.
APA
Sherstan, C., Ashley, D.R., Bennett, B., Young, K., White, A., White, M. & Sutton, R.S.. (2018). Comparing Direct and Indirect Temporal-Difference Methods for Estimating the Variance of the Return. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R16:62-71 Available from https://proceedings.mlr.press/r16/sherstan18a.html. Reissued by PMLR on 04 October 2026.

Related Material