Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments

Ege Can Kaya, Mahsa Ghasemi, Abolfazl Hashemi
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2877-2893, 2026.

Abstract

Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority. However, the classical {Markov} decision process ({MDP}) formalism specifies only marginal laws and leaves the joint law of counterfactual one-step outcomes across multiple possible actions at a state unspecified. We study coupled-dynamics environments with a multi-action generative interface which can sample counterfactual one-step outcomes for multiple actions under shared exogenous randomness. We propose joint {MDPs} (JMDPs) as a formalism for such environments by augmenting an {MDP} with a multi-action sample transition model which specifies a coupling of one-step counterfactual outcomes, while preserving standard {MDP} interaction as marginal observations. We adopt and formalize a one-step coupling regime where dependence across actions is confined to immediate counterfactual outcomes at the queried state. In this regime, we derive {Bellman} operators for $n$th-order return moments, providing dynamic programming and incremental algorithms with convergence guarantees.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kaya26b, title = {Joint {MDPs} and Reinforcement Learning in Coupled-Dynamics Environments}, author = {Kaya, Ege Can and Ghasemi, Mahsa and Hashemi, Abolfazl}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2877--2893}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kaya26b/kaya26b.pdf}, url = {https://proceedings.mlr.press/v337/kaya26b.html}, abstract = {Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority. However, the classical {Markov} decision process ({MDP}) formalism specifies only marginal laws and leaves the joint law of counterfactual one-step outcomes across multiple possible actions at a state unspecified. We study coupled-dynamics environments with a multi-action generative interface which can sample counterfactual one-step outcomes for multiple actions under shared exogenous randomness. We propose joint {MDPs} (JMDPs) as a formalism for such environments by augmenting an {MDP} with a multi-action sample transition model which specifies a coupling of one-step counterfactual outcomes, while preserving standard {MDP} interaction as marginal observations. We adopt and formalize a one-step coupling regime where dependence across actions is confined to immediate counterfactual outcomes at the queried state. In this regime, we derive {Bellman} operators for $n$th-order return moments, providing dynamic programming and incremental algorithms with convergence guarantees.} }
Endnote
%0 Conference Paper %T Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments %A Ege Can Kaya %A Mahsa Ghasemi %A Abolfazl Hashemi %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kaya26b %I PMLR %P 2877--2893 %U https://proceedings.mlr.press/v337/kaya26b.html %V 337 %X Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority. However, the classical {Markov} decision process ({MDP}) formalism specifies only marginal laws and leaves the joint law of counterfactual one-step outcomes across multiple possible actions at a state unspecified. We study coupled-dynamics environments with a multi-action generative interface which can sample counterfactual one-step outcomes for multiple actions under shared exogenous randomness. We propose joint {MDPs} (JMDPs) as a formalism for such environments by augmenting an {MDP} with a multi-action sample transition model which specifies a coupling of one-step counterfactual outcomes, while preserving standard {MDP} interaction as marginal observations. We adopt and formalize a one-step coupling regime where dependence across actions is confined to immediate counterfactual outcomes at the queried state. In this regime, we derive {Bellman} operators for $n$th-order return moments, providing dynamic programming and incremental algorithms with convergence guarantees.
APA
Kaya, E.C., Ghasemi, M. & Hashemi, A.. (2026). Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2877-2893 Available from https://proceedings.mlr.press/v337/kaya26b.html.

Related Material