Multivariate Distributional Reinforcement Learning Using Sliced Divergences

Baptiste Debes, Tinne Tuytelaars
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:23456-23526, 2026.

Abstract

Distributional reinforcement learning (DRL) models the full return distribution rather than expectations, but extending it to multivariate settings remains challenging. Many common metrics do not naturally generalize beyond one dimension or lose computational tractability, and the multivariate case introduces additional difficulties such as general matrix discounting, for which no contraction results are available. We introduce Sliced Distributional Reinforcement Learning (SDRL), which lifts tractable one-dimensional divergences to multivariate return distributions via projections. We prove Bellman contraction for uniform slicing under shared scalar discounting, and introduce a maximum-slicing variant with contraction under general dense discount matrices. SDRL supports a broad class of base divergences; we analyze Wasserstein, Cramér, and Maximum Mean Discrepancy (MMD), and characterize which SDRL variants suit the standard single-sample Bellman update used in distributional RL. We evaluate SDRL on a toy chain problem and a gridworld image-based environment as well as a subset of Atari games. Code is available at https://github.com/BaptisteDebes/SlicedDistributionalRL

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-debes26a, title = {Multivariate Distributional Reinforcement Learning Using Sliced Divergences}, author = {Debes, Baptiste and Tuytelaars, Tinne}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {23456--23526}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/debes26a/debes26a.pdf}, url = {https://proceedings.mlr.press/v306/debes26a.html}, abstract = {Distributional reinforcement learning (DRL) models the full return distribution rather than expectations, but extending it to multivariate settings remains challenging. Many common metrics do not naturally generalize beyond one dimension or lose computational tractability, and the multivariate case introduces additional difficulties such as general matrix discounting, for which no contraction results are available. We introduce Sliced Distributional Reinforcement Learning (SDRL), which lifts tractable one-dimensional divergences to multivariate return distributions via projections. We prove Bellman contraction for uniform slicing under shared scalar discounting, and introduce a maximum-slicing variant with contraction under general dense discount matrices. SDRL supports a broad class of base divergences; we analyze Wasserstein, Cramér, and Maximum Mean Discrepancy (MMD), and characterize which SDRL variants suit the standard single-sample Bellman update used in distributional RL. We evaluate SDRL on a toy chain problem and a gridworld image-based environment as well as a subset of Atari games. Code is available at https://github.com/BaptisteDebes/SlicedDistributionalRL} }
Endnote
%0 Conference Paper %T Multivariate Distributional Reinforcement Learning Using Sliced Divergences %A Baptiste Debes %A Tinne Tuytelaars %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-debes26a %I PMLR %P 23456--23526 %U https://proceedings.mlr.press/v306/debes26a.html %V 306 %X Distributional reinforcement learning (DRL) models the full return distribution rather than expectations, but extending it to multivariate settings remains challenging. Many common metrics do not naturally generalize beyond one dimension or lose computational tractability, and the multivariate case introduces additional difficulties such as general matrix discounting, for which no contraction results are available. We introduce Sliced Distributional Reinforcement Learning (SDRL), which lifts tractable one-dimensional divergences to multivariate return distributions via projections. We prove Bellman contraction for uniform slicing under shared scalar discounting, and introduce a maximum-slicing variant with contraction under general dense discount matrices. SDRL supports a broad class of base divergences; we analyze Wasserstein, Cramér, and Maximum Mean Discrepancy (MMD), and characterize which SDRL variants suit the standard single-sample Bellman update used in distributional RL. We evaluate SDRL on a toy chain problem and a gridworld image-based environment as well as a subset of Atari games. Code is available at https://github.com/BaptisteDebes/SlicedDistributionalRL
APA
Debes, B. & Tuytelaars, T.. (2026). Multivariate Distributional Reinforcement Learning Using Sliced Divergences. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:23456-23526 Available from https://proceedings.mlr.press/v306/debes26a.html.

Related Material