Noise as a Natural Regularizer in Markov Decision Processes: Connecting Environmental Stochasticity and Policy Simplicity

Harry Chen, Yiyang Sun, Michal Moshkovitz, Zachery Boner, Lesia Semenova, Cynthia Rudin, Ronald Parr
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16742-16767, 2026.

Abstract

The planning horizon in a Markov Decision Process (MDP) determines how far into the future an agent reasons. In practice, shorter horizons are commonly associated with policies that exhibit simpler or more interpretable decision-making behavior. In this paper, we establish a formal connection between environmental stochasticity and planning horizon in MDPs. We show that for broad classes of transition noise, solving a noisy MDP can be formally related to solving a noise-free MDP with a shorter effective discount factor, leading to identical optimal policies in some cases and near-optimal ones in others. We further characterize settings in which this correspondence breaks down, clarifying when horizon-based interpretations of noise are not valid. These results, which are supported by both theory and experiments, also give some insight into the common practice of using smaller discount factors for reinforcement learning than those that can be justified by standard modeling interpretations.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26dy, title = {Noise as a Natural Regularizer in {M}arkov Decision Processes: Connecting Environmental Stochasticity and Policy Simplicity}, author = {Chen, Harry and Sun, Yiyang and Moshkovitz, Michal and Boner, Zachery and Semenova, Lesia and Rudin, Cynthia and Parr, Ronald}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16742--16767}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26dy/chen26dy.pdf}, url = {https://proceedings.mlr.press/v306/chen26dy.html}, abstract = {The planning horizon in a Markov Decision Process (MDP) determines how far into the future an agent reasons. In practice, shorter horizons are commonly associated with policies that exhibit simpler or more interpretable decision-making behavior. In this paper, we establish a formal connection between environmental stochasticity and planning horizon in MDPs. We show that for broad classes of transition noise, solving a noisy MDP can be formally related to solving a noise-free MDP with a shorter effective discount factor, leading to identical optimal policies in some cases and near-optimal ones in others. We further characterize settings in which this correspondence breaks down, clarifying when horizon-based interpretations of noise are not valid. These results, which are supported by both theory and experiments, also give some insight into the common practice of using smaller discount factors for reinforcement learning than those that can be justified by standard modeling interpretations.} }
Endnote
%0 Conference Paper %T Noise as a Natural Regularizer in Markov Decision Processes: Connecting Environmental Stochasticity and Policy Simplicity %A Harry Chen %A Yiyang Sun %A Michal Moshkovitz %A Zachery Boner %A Lesia Semenova %A Cynthia Rudin %A Ronald Parr %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26dy %I PMLR %P 16742--16767 %U https://proceedings.mlr.press/v306/chen26dy.html %V 306 %X The planning horizon in a Markov Decision Process (MDP) determines how far into the future an agent reasons. In practice, shorter horizons are commonly associated with policies that exhibit simpler or more interpretable decision-making behavior. In this paper, we establish a formal connection between environmental stochasticity and planning horizon in MDPs. We show that for broad classes of transition noise, solving a noisy MDP can be formally related to solving a noise-free MDP with a shorter effective discount factor, leading to identical optimal policies in some cases and near-optimal ones in others. We further characterize settings in which this correspondence breaks down, clarifying when horizon-based interpretations of noise are not valid. These results, which are supported by both theory and experiments, also give some insight into the common practice of using smaller discount factors for reinforcement learning than those that can be justified by standard modeling interpretations.
APA
Chen, H., Sun, Y., Moshkovitz, M., Boner, Z., Semenova, L., Rudin, C. & Parr, R.. (2026). Noise as a Natural Regularizer in Markov Decision Processes: Connecting Environmental Stochasticity and Policy Simplicity. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16742-16767 Available from https://proceedings.mlr.press/v306/chen26dy.html.

Related Material