Efficient Learning of Deep State Space Models via Importance Smoothing

John-Joseph Brady, Nikolas Nüsken, Yunpeng Li
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:9566-9597, 2026.

Abstract

Latent state space systems are ubiquitous in statistical modelling, arising naturally when time series are observed through noisy measurements. However, training deep state space models (DSSMs) at scale remains difficult. Two largely distinct strategies have emerged for training DSSMs. The first, auto-encoding DSSMs, trains generative models by optimising a variational lower bound. The second backpropagates through the outputs of classical sequential Monte Carlo (SMC) algorithms. Such approaches can train DSSMs for both discriminative and generative tasks, but their inherently sequential forward passes scale poorly on modern hardware. We propose parallel variational Monte Carlo (PVMC), a new training method that bridges these paradigms and robustly trains DSSMs for both discriminative and generative tasks. Across a set of benchmark experiments, PVMC matches or exceeds state-of-the-art performance while training $10\times$ faster than the fastest competing SMC-based approach.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-brady26a, title = {Efficient Learning of Deep State Space Models via Importance Smoothing}, author = {Brady, John-Joseph and N\"{u}sken, Nikolas and Li, Yunpeng}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {9566--9597}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/brady26a/brady26a.pdf}, url = {https://proceedings.mlr.press/v306/brady26a.html}, abstract = {Latent state space systems are ubiquitous in statistical modelling, arising naturally when time series are observed through noisy measurements. However, training deep state space models (DSSMs) at scale remains difficult. Two largely distinct strategies have emerged for training DSSMs. The first, auto-encoding DSSMs, trains generative models by optimising a variational lower bound. The second backpropagates through the outputs of classical sequential Monte Carlo (SMC) algorithms. Such approaches can train DSSMs for both discriminative and generative tasks, but their inherently sequential forward passes scale poorly on modern hardware. We propose parallel variational Monte Carlo (PVMC), a new training method that bridges these paradigms and robustly trains DSSMs for both discriminative and generative tasks. Across a set of benchmark experiments, PVMC matches or exceeds state-of-the-art performance while training $10\times$ faster than the fastest competing SMC-based approach.} }
Endnote
%0 Conference Paper %T Efficient Learning of Deep State Space Models via Importance Smoothing %A John-Joseph Brady %A Nikolas Nüsken %A Yunpeng Li %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-brady26a %I PMLR %P 9566--9597 %U https://proceedings.mlr.press/v306/brady26a.html %V 306 %X Latent state space systems are ubiquitous in statistical modelling, arising naturally when time series are observed through noisy measurements. However, training deep state space models (DSSMs) at scale remains difficult. Two largely distinct strategies have emerged for training DSSMs. The first, auto-encoding DSSMs, trains generative models by optimising a variational lower bound. The second backpropagates through the outputs of classical sequential Monte Carlo (SMC) algorithms. Such approaches can train DSSMs for both discriminative and generative tasks, but their inherently sequential forward passes scale poorly on modern hardware. We propose parallel variational Monte Carlo (PVMC), a new training method that bridges these paradigms and robustly trains DSSMs for both discriminative and generative tasks. Across a set of benchmark experiments, PVMC matches or exceeds state-of-the-art performance while training $10\times$ faster than the fastest competing SMC-based approach.
APA
Brady, J., Nüsken, N. & Li, Y.. (2026). Efficient Learning of Deep State Space Models via Importance Smoothing. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:9566-9597 Available from https://proceedings.mlr.press/v306/brady26a.html.

Related Material