[edit]
Comparing Direct and Indirect Temporal-Difference Methods for Estimating the Variance of the Return
Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence, PMLR R16:62-71, 2018.
Abstract
Temporal-difference (TD) learning methods are widely used in reinforcement learning to estimate the expected return for each state, without a model, because of their significant advantages in computational and data effi- ciency. For many applications involving risk mitigation, it would also be useful to estimate the variance of the return by TD methods. In this paper, we describe a way of doing this that is substantially simpler than those proposed by Tamar, Di Castro, and Mannor in 2012, or those proposed by White and White in 2016. We show that two TD learners operating in series can learn expectation and variance esti- mates. The trick is to use the square of the TD error of the expectation learner as the reward of the variance learner, and the square of the ex- pectation learner’s discount rate as the discount rate of the variance learner. With these two modifications, the variance learning problem becomes a conventional TD learning problem to which standard theoretical results can be ap- plied. Our formal results are limited to the ta- ble lookup case, for which our method is still novel, but the extension to function approxi- mation is immediate, and we provide some em- pirical results for the linear function approx- imation case. Our experimental results show that our direct method behaves just as well as a comparable indirect method, but is generally more robust.