An NTK Theory Approach to UCB Estimation in Semi-gradient TD Learning

Yijun Wu, Pascal R. van der Vaart, Moritz Akiya Zanger, Matthijs T. J. Spaan
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:7403-7415, 2026.

Abstract

Efficient exploration is a central challenge in reinforcement learning ({RL}), especially under sparse rewards, large state-action spaces, or adversarial environment dynamics. A principled approach to exploration is to choose actions that maximize the upper bound return. Under certain conditions, this can be formulated in the {Bayesian} framework as exploration based on the posterior variance. In practice, the predictive variance of a neural network is often used as a surrogate for the posterior variance by techniques such as ensembles, randomized priors, and random network distillation. However, for wide neural networks, neural tangent kernel ({NTK}) theory proves there is a discrepancy between the predictive variance of a wide neural network trained by gradient descent and its posterior variance. This is a known issue in supervised learning and there exist methods to reconcile this discrepancy. In this work, we show that, in contrast to the supervised setting, reconciliation is not possible for semi-gradient temporal difference ({TD}) learning, which is a commonly used algorithm in deep {RL}. Instead, we derive the upper confidence bound ({UCB}) of the value function in semi-gradient {TD} learning directly and show that this {UCB} can be estimated by a neural network. To aid our analysis, we extend {NTK} theory to regularized semi-gradient {TD} learning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-wu26a, title = {An {NTK} Theory Approach to {UCB} Estimation in Semi-gradient {TD} Learning}, author = {Wu, Yijun and van der Vaart, Pascal R. and Zanger, Moritz Akiya and Spaan, Matthijs T. J.}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {7403--7415}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/wu26a/wu26a.pdf}, url = {https://proceedings.mlr.press/v337/wu26a.html}, abstract = {Efficient exploration is a central challenge in reinforcement learning ({RL}), especially under sparse rewards, large state-action spaces, or adversarial environment dynamics. A principled approach to exploration is to choose actions that maximize the upper bound return. Under certain conditions, this can be formulated in the {Bayesian} framework as exploration based on the posterior variance. In practice, the predictive variance of a neural network is often used as a surrogate for the posterior variance by techniques such as ensembles, randomized priors, and random network distillation. However, for wide neural networks, neural tangent kernel ({NTK}) theory proves there is a discrepancy between the predictive variance of a wide neural network trained by gradient descent and its posterior variance. This is a known issue in supervised learning and there exist methods to reconcile this discrepancy. In this work, we show that, in contrast to the supervised setting, reconciliation is not possible for semi-gradient temporal difference ({TD}) learning, which is a commonly used algorithm in deep {RL}. Instead, we derive the upper confidence bound ({UCB}) of the value function in semi-gradient {TD} learning directly and show that this {UCB} can be estimated by a neural network. To aid our analysis, we extend {NTK} theory to regularized semi-gradient {TD} learning.} }
Endnote
%0 Conference Paper %T An NTK Theory Approach to UCB Estimation in Semi-gradient TD Learning %A Yijun Wu %A Pascal R. van der Vaart %A Moritz Akiya Zanger %A Matthijs T. J. Spaan %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-wu26a %I PMLR %P 7403--7415 %U https://proceedings.mlr.press/v337/wu26a.html %V 337 %X Efficient exploration is a central challenge in reinforcement learning ({RL}), especially under sparse rewards, large state-action spaces, or adversarial environment dynamics. A principled approach to exploration is to choose actions that maximize the upper bound return. Under certain conditions, this can be formulated in the {Bayesian} framework as exploration based on the posterior variance. In practice, the predictive variance of a neural network is often used as a surrogate for the posterior variance by techniques such as ensembles, randomized priors, and random network distillation. However, for wide neural networks, neural tangent kernel ({NTK}) theory proves there is a discrepancy between the predictive variance of a wide neural network trained by gradient descent and its posterior variance. This is a known issue in supervised learning and there exist methods to reconcile this discrepancy. In this work, we show that, in contrast to the supervised setting, reconciliation is not possible for semi-gradient temporal difference ({TD}) learning, which is a commonly used algorithm in deep {RL}. Instead, we derive the upper confidence bound ({UCB}) of the value function in semi-gradient {TD} learning directly and show that this {UCB} can be estimated by a neural network. To aid our analysis, we extend {NTK} theory to regularized semi-gradient {TD} learning.
APA
Wu, Y., van der Vaart, P.R., Zanger, M.A. & Spaan, M.T.J.. (2026). An NTK Theory Approach to UCB Estimation in Semi-gradient TD Learning. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:7403-7415 Available from https://proceedings.mlr.press/v337/wu26a.html.

Related Material