Taming the Noise in Reinforcement Learning via Soft Updates

Roy Fox HUJI, Ari Pakman Columbia University, Naftali Tishby HUJI
Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, PMLR R14:582-591, 2016.

Abstract

Model-free reinforcement learning algorithms, such as Q-learning, perform poorly in the early stages of learning in noisy environments, because much effort is spent unlearning biased estimates of the state-action value function. The bias results from selecting, among several noisy estimates, the apparent optimum, which may actually be suboptimal. We propose G-learning, a new off-policy learning algorithm that regularizes the value estimates by penalizing deterministic policies in the beginning of the learning process. We show that this method reduces the bias of the value-function estimation, leading to faster convergence to the optimal value and the optimal policy. Moreover, G-learning enables the natural incorporation of prior domain knowledge, when available. The stochastic nature of G-learning also makes it avoid some exploration costs, a property usually attributed only to on-policy algorithms. We illustrate these ideas in several examples, where G-learning results in significant improvements of the convergence rate and the cost of the learning process.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR14-huji16a, title = {Taming the Noise in Reinforcement Learning via Soft Updates}, author = {HUJI, Roy Fox and University, Ari Pakman Columbia and HUJI, Naftali Tishby}, booktitle = {Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence}, pages = {582--591}, year = {2016}, editor = {Ihler, Alexander and Janzing, Dominik}, volume = {R14}, series = {Proceedings of Machine Learning Research}, month = {25--29 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r14/main/assets/huji16a/huji16a.pdf}, url = {https://proceedings.mlr.press/r14/huji16a.html}, abstract = {Model-free reinforcement learning algorithms, such as Q-learning, perform poorly in the early stages of learning in noisy environments, because much effort is spent unlearning biased estimates of the state-action value function. The bias results from selecting, among several noisy estimates, the apparent optimum, which may actually be suboptimal. We propose G-learning, a new off-policy learning algorithm that regularizes the value estimates by penalizing deterministic policies in the beginning of the learning process. We show that this method reduces the bias of the value-function estimation, leading to faster convergence to the optimal value and the optimal policy. Moreover, G-learning enables the natural incorporation of prior domain knowledge, when available. The stochastic nature of G-learning also makes it avoid some exploration costs, a property usually attributed only to on-policy algorithms. We illustrate these ideas in several examples, where G-learning results in significant improvements of the convergence rate and the cost of the learning process.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Taming the Noise in Reinforcement Learning via Soft Updates %A Roy Fox HUJI %A Ari Pakman Columbia University %A Naftali Tishby HUJI %B Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2016 %E Alexander Ihler %E Dominik Janzing %F pmlr-vR14-huji16a %I PMLR %P 582--591 %U https://proceedings.mlr.press/r14/huji16a.html %V R14 %X Model-free reinforcement learning algorithms, such as Q-learning, perform poorly in the early stages of learning in noisy environments, because much effort is spent unlearning biased estimates of the state-action value function. The bias results from selecting, among several noisy estimates, the apparent optimum, which may actually be suboptimal. We propose G-learning, a new off-policy learning algorithm that regularizes the value estimates by penalizing deterministic policies in the beginning of the learning process. We show that this method reduces the bias of the value-function estimation, leading to faster convergence to the optimal value and the optimal policy. Moreover, G-learning enables the natural incorporation of prior domain knowledge, when available. The stochastic nature of G-learning also makes it avoid some exploration costs, a property usually attributed only to on-policy algorithms. We illustrate these ideas in several examples, where G-learning results in significant improvements of the convergence rate and the cost of the learning process. %Z Reissued by PMLR on 04 October 2026.
APA
HUJI, R.F., University, A.P.C. & HUJI, N.T.. (2016). Taming the Noise in Reinforcement Learning via Soft Updates. Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R14:582-591 Available from https://proceedings.mlr.press/r14/huji16a.html. Reissued by PMLR on 04 October 2026.

Related Material