Sparse Action-Dependent Policy Iteration under Coordination Structures: Convergence and Optimality

Jianglin Ding, Jingcheng Tang, Gangshan Jing
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1423-1447, 2026.

Abstract

Action-dependent policies, which condition decisions of each agent on both states and other agents’ actions, provide a powerful structured framework for cooperative multi-agent reinforcement learning ({MARL}). Most existing studies have focused on auto-regressive formulations, where each agent’s policy depends on the actions of all preceding agents. However, this structure suffers from severe scalability limitations as the number of agents grows. In contrast, the theoretical foundations of sparse dependency structures remain largely unexplored. To address this gap, we introduce the Action Dependency Graph ({ADG}) to model sparse inter-agent dependencies. We propose a refined equilibrium concept with respect to the {ADG} that is stronger than the {Nash} equilibrium which often traps independent policies. Furthermore, within Coordination Graphs (CG) structured problems, we show that such an equilibrium attains global optimality when the {ADG} satisfies specific CG-induced conditions. To substantiate our theory, we develop a tabular multi-agent policy iteration algorithm that converges to the refined equilibrium exactly as predicted. We further extend our approach to deep {MARL}, confirming that these structural conditions provide a reliable design principle for scalable and optimal coordination.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-ding26a, title = {Sparse Action-Dependent Policy Iteration under Coordination Structures: Convergence and Optimality}, author = {Ding, Jianglin and Tang, Jingcheng and Jing, Gangshan}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {1423--1447}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/ding26a/ding26a.pdf}, url = {https://proceedings.mlr.press/v337/ding26a.html}, abstract = {Action-dependent policies, which condition decisions of each agent on both states and other agents’ actions, provide a powerful structured framework for cooperative multi-agent reinforcement learning ({MARL}). Most existing studies have focused on auto-regressive formulations, where each agent’s policy depends on the actions of all preceding agents. However, this structure suffers from severe scalability limitations as the number of agents grows. In contrast, the theoretical foundations of sparse dependency structures remain largely unexplored. To address this gap, we introduce the Action Dependency Graph ({ADG}) to model sparse inter-agent dependencies. We propose a refined equilibrium concept with respect to the {ADG} that is stronger than the {Nash} equilibrium which often traps independent policies. Furthermore, within Coordination Graphs (CG) structured problems, we show that such an equilibrium attains global optimality when the {ADG} satisfies specific CG-induced conditions. To substantiate our theory, we develop a tabular multi-agent policy iteration algorithm that converges to the refined equilibrium exactly as predicted. We further extend our approach to deep {MARL}, confirming that these structural conditions provide a reliable design principle for scalable and optimal coordination.} }
Endnote
%0 Conference Paper %T Sparse Action-Dependent Policy Iteration under Coordination Structures: Convergence and Optimality %A Jianglin Ding %A Jingcheng Tang %A Gangshan Jing %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-ding26a %I PMLR %P 1423--1447 %U https://proceedings.mlr.press/v337/ding26a.html %V 337 %X Action-dependent policies, which condition decisions of each agent on both states and other agents’ actions, provide a powerful structured framework for cooperative multi-agent reinforcement learning ({MARL}). Most existing studies have focused on auto-regressive formulations, where each agent’s policy depends on the actions of all preceding agents. However, this structure suffers from severe scalability limitations as the number of agents grows. In contrast, the theoretical foundations of sparse dependency structures remain largely unexplored. To address this gap, we introduce the Action Dependency Graph ({ADG}) to model sparse inter-agent dependencies. We propose a refined equilibrium concept with respect to the {ADG} that is stronger than the {Nash} equilibrium which often traps independent policies. Furthermore, within Coordination Graphs (CG) structured problems, we show that such an equilibrium attains global optimality when the {ADG} satisfies specific CG-induced conditions. To substantiate our theory, we develop a tabular multi-agent policy iteration algorithm that converges to the refined equilibrium exactly as predicted. We further extend our approach to deep {MARL}, confirming that these structural conditions provide a reliable design principle for scalable and optimal coordination.
APA
Ding, J., Tang, J. & Jing, G.. (2026). Sparse Action-Dependent Policy Iteration under Coordination Structures: Convergence and Optimality. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:1423-1447 Available from https://proceedings.mlr.press/v337/ding26a.html.

Related Material