Off-policy TD($ł$) with a true online equivalence

Hado Van Hasselt, Rupam Mahmood, Rich Sutton
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:882-891, 2014.

Abstract

Van Seijen and Sutton (2014) recently proposed a new version of the linear TD($\lambda$) learning algo- rithm that is exactly equivalent to an online for- ward view and that empirically performed bet- ter than its classical counterpart in both predic- tion and control problems. However, their al- gorithm is restricted to on-policy learning. In the more general case of off-policy learning, in which the policy whose outcome is predicted and the policy used to generate data may be differ- ent, their algorithm cannot be applied. One rea- son for this is that the algorithm bootstraps and thus is subject to instability problems when func- tion approximation is used. A second reason true online TD($\lambda$) cannot be used for off-policy learning is that the off-policy case requires so- phisticated importance sampling in its eligibility traces. To address these limitations, we gener- alize their equivalence result and use this gen- eralization to construct the first online algorithm to be exactly equivalent to an off-policy forward view. We show this algorithm, named true on- line GTD($\lambda$), empirically outperforms GTD($\lambda$) (Maei, 2011) which was derived from the same objective as our forward view but lacks the ex- act online equivalence. In the general theorem that allows us to derive this new algorithm, we encounter a new general eligibility-trace update.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-van-hasselt14a, title = {Off-policy {TD}($ł$) with a true online equivalence}, author = {Van Hasselt, Hado and Mahmood, Rupam and Sutton, Rich}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {882--891}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/van-hasselt14a/van-hasselt14a.pdf}, url = {https://proceedings.mlr.press/r12/van-hasselt14a.html}, abstract = {Van Seijen and Sutton (2014) recently proposed a new version of the linear TD($\lambda$) learning algo- rithm that is exactly equivalent to an online for- ward view and that empirically performed bet- ter than its classical counterpart in both predic- tion and control problems. However, their al- gorithm is restricted to on-policy learning. In the more general case of off-policy learning, in which the policy whose outcome is predicted and the policy used to generate data may be differ- ent, their algorithm cannot be applied. One rea- son for this is that the algorithm bootstraps and thus is subject to instability problems when func- tion approximation is used. A second reason true online TD($\lambda$) cannot be used for off-policy learning is that the off-policy case requires so- phisticated importance sampling in its eligibility traces. To address these limitations, we gener- alize their equivalence result and use this gen- eralization to construct the first online algorithm to be exactly equivalent to an off-policy forward view. We show this algorithm, named true on- line GTD($\lambda$), empirically outperforms GTD($\lambda$) (Maei, 2011) which was derived from the same objective as our forward view but lacks the ex- act online equivalence. In the general theorem that allows us to derive this new algorithm, we encounter a new general eligibility-trace update.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Off-policy TD($ł$) with a true online equivalence %A Hado Van Hasselt %A Rupam Mahmood %A Rich Sutton %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-van-hasselt14a %I PMLR %P 882--891 %U https://proceedings.mlr.press/r12/van-hasselt14a.html %V R12 %X Van Seijen and Sutton (2014) recently proposed a new version of the linear TD($\lambda$) learning algo- rithm that is exactly equivalent to an online for- ward view and that empirically performed bet- ter than its classical counterpart in both predic- tion and control problems. However, their al- gorithm is restricted to on-policy learning. In the more general case of off-policy learning, in which the policy whose outcome is predicted and the policy used to generate data may be differ- ent, their algorithm cannot be applied. One rea- son for this is that the algorithm bootstraps and thus is subject to instability problems when func- tion approximation is used. A second reason true online TD($\lambda$) cannot be used for off-policy learning is that the off-policy case requires so- phisticated importance sampling in its eligibility traces. To address these limitations, we gener- alize their equivalence result and use this gen- eralization to construct the first online algorithm to be exactly equivalent to an off-policy forward view. We show this algorithm, named true on- line GTD($\lambda$), empirically outperforms GTD($\lambda$) (Maei, 2011) which was derived from the same objective as our forward view but lacks the ex- act online equivalence. In the general theorem that allows us to derive this new algorithm, we encounter a new general eligibility-trace update. %Z Reissued by PMLR on 04 October 2026.
APA
Van Hasselt, H., Mahmood, R. & Sutton, R.. (2014). Off-policy TD($ł$) with a true online equivalence. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:882-891 Available from https://proceedings.mlr.press/r12/van-hasselt14a.html. Reissued by PMLR on 04 October 2026.

Related Material