[edit]
Sparse Action-Dependent Policy Iteration under Coordination Structures: Convergence and Optimality
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1423-1447, 2026.
Abstract
Action-dependent policies, which condition decisions of each agent on both states and other agents’ actions, provide a powerful structured framework for cooperative multi-agent reinforcement learning ({MARL}). Most existing studies have focused on auto-regressive formulations, where each agent’s policy depends on the actions of all preceding agents. However, this structure suffers from severe scalability limitations as the number of agents grows. In contrast, the theoretical foundations of sparse dependency structures remain largely unexplored. To address this gap, we introduce the Action Dependency Graph ({ADG}) to model sparse inter-agent dependencies. We propose a refined equilibrium concept with respect to the {ADG} that is stronger than the {Nash} equilibrium which often traps independent policies. Furthermore, within Coordination Graphs (CG) structured problems, we show that such an equilibrium attains global optimality when the {ADG} satisfies specific CG-induced conditions. To substantiate our theory, we develop a tabular multi-agent policy iteration algorithm that converges to the refined equilibrium exactly as predicted. We further extend our approach to deep {MARL}, confirming that these structural conditions provide a reliable design principle for scalable and optimal coordination.