Error Propagation in Dynamic Programming: From Stochastic Control to American Option Pricing

Andrea Della Vecchia, Damir Filipovic
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:23637-23652, 2026.

Abstract

This paper investigates theoretical and methodological foundations for stochastic optimal control (SOC) in discrete time. We start formulating the control problem in a general dynamic programming framework, introducing the mathematical structure needed for a detailed convergence analysis. The associate value function is estimated through a sequence of approximations combining nonparametric regression methods and Monte Carlo subsampling. The regression step is performed within reproducing kernel Hilbert spaces (RKHSs), exploiting the classical KRR algorithm, while Monte Carlo sampling methods are introduced to estimate the continuation value. To assess the accuracy of our value function estimator, we propose a natural error decomposition and rigorously control the resulting error terms at each time step. We then analyze how this error propagates backward in time-from maturity to the initial stage-a relatively underexplored aspect of the SOC literature. Finally, we illustrate how our analysis naturally applies to a key financial application: the pricing of American options.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-della-vecchia26a, title = {Error Propagation in Dynamic Programming: From Stochastic Control to American Option Pricing}, author = {Della Vecchia, Andrea and Filipovic, Damir}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {23637--23652}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/della-vecchia26a/della-vecchia26a.pdf}, url = {https://proceedings.mlr.press/v306/della-vecchia26a.html}, abstract = {This paper investigates theoretical and methodological foundations for stochastic optimal control (SOC) in discrete time. We start formulating the control problem in a general dynamic programming framework, introducing the mathematical structure needed for a detailed convergence analysis. The associate value function is estimated through a sequence of approximations combining nonparametric regression methods and Monte Carlo subsampling. The regression step is performed within reproducing kernel Hilbert spaces (RKHSs), exploiting the classical KRR algorithm, while Monte Carlo sampling methods are introduced to estimate the continuation value. To assess the accuracy of our value function estimator, we propose a natural error decomposition and rigorously control the resulting error terms at each time step. We then analyze how this error propagates backward in time-from maturity to the initial stage-a relatively underexplored aspect of the SOC literature. Finally, we illustrate how our analysis naturally applies to a key financial application: the pricing of American options.} }
Endnote
%0 Conference Paper %T Error Propagation in Dynamic Programming: From Stochastic Control to American Option Pricing %A Andrea Della Vecchia %A Damir Filipovic %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-della-vecchia26a %I PMLR %P 23637--23652 %U https://proceedings.mlr.press/v306/della-vecchia26a.html %V 306 %X This paper investigates theoretical and methodological foundations for stochastic optimal control (SOC) in discrete time. We start formulating the control problem in a general dynamic programming framework, introducing the mathematical structure needed for a detailed convergence analysis. The associate value function is estimated through a sequence of approximations combining nonparametric regression methods and Monte Carlo subsampling. The regression step is performed within reproducing kernel Hilbert spaces (RKHSs), exploiting the classical KRR algorithm, while Monte Carlo sampling methods are introduced to estimate the continuation value. To assess the accuracy of our value function estimator, we propose a natural error decomposition and rigorously control the resulting error terms at each time step. We then analyze how this error propagates backward in time-from maturity to the initial stage-a relatively underexplored aspect of the SOC literature. Finally, we illustrate how our analysis naturally applies to a key financial application: the pricing of American options.
APA
Della Vecchia, A. & Filipovic, D.. (2026). Error Propagation in Dynamic Programming: From Stochastic Control to American Option Pricing. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:23637-23652 Available from https://proceedings.mlr.press/v306/della-vecchia26a.html.

Related Material