Multi-Agent RL with Invisible Collaborators: Marginal Advantage Estimation for Indirect Cooperation

Jianglin Qiao, Zehong Cao, Siyi Hu, Mingjun Fan, Mahardhika Pratama, Ryszard Kowalczyk
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:5573-5600, 2026.

Abstract

Cooperative Multi-Agent Reinforcement Learning ({MARL}) has primarily focused on direct cooperation among agents. However, many real-world systems exhibit structurally asymmetric cooperation, where agents are physically constrained and cannot directly coordinate with those on whom they depend. In such settings, accurately estimating individual contributions under high uncertainty is challenging. We propose Marginal Advantage Estimation (MAE), a representation-level contribution estimation method for cooperative {MARL} under this structural isolation. MAE employs a synchronised feature-masking mechanism to evaluate marginal contributions without action-level counterfactual perturbations, thereby reducing variance and providing more informative learning signals. We provide a theoretical analysis of its bias–variance properties and demonstrate consistent performance improvements across 18 tasks in Partitioned MPE and Basilisk benchmarks over strong cooperative {MARL} baselines.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-qiao26a, title = {Multi-Agent {RL} with Invisible Collaborators: Marginal Advantage Estimation for Indirect Cooperation}, author = {Qiao, Jianglin and Cao, Zehong and Hu, Siyi and Fan, Mingjun and Pratama, Mahardhika and Kowalczyk, Ryszard}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {5573--5600}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/qiao26a/qiao26a.pdf}, url = {https://proceedings.mlr.press/v337/qiao26a.html}, abstract = {Cooperative Multi-Agent Reinforcement Learning ({MARL}) has primarily focused on direct cooperation among agents. However, many real-world systems exhibit structurally asymmetric cooperation, where agents are physically constrained and cannot directly coordinate with those on whom they depend. In such settings, accurately estimating individual contributions under high uncertainty is challenging. We propose Marginal Advantage Estimation (MAE), a representation-level contribution estimation method for cooperative {MARL} under this structural isolation. MAE employs a synchronised feature-masking mechanism to evaluate marginal contributions without action-level counterfactual perturbations, thereby reducing variance and providing more informative learning signals. We provide a theoretical analysis of its bias–variance properties and demonstrate consistent performance improvements across 18 tasks in Partitioned MPE and Basilisk benchmarks over strong cooperative {MARL} baselines.} }
Endnote
%0 Conference Paper %T Multi-Agent RL with Invisible Collaborators: Marginal Advantage Estimation for Indirect Cooperation %A Jianglin Qiao %A Zehong Cao %A Siyi Hu %A Mingjun Fan %A Mahardhika Pratama %A Ryszard Kowalczyk %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-qiao26a %I PMLR %P 5573--5600 %U https://proceedings.mlr.press/v337/qiao26a.html %V 337 %X Cooperative Multi-Agent Reinforcement Learning ({MARL}) has primarily focused on direct cooperation among agents. However, many real-world systems exhibit structurally asymmetric cooperation, where agents are physically constrained and cannot directly coordinate with those on whom they depend. In such settings, accurately estimating individual contributions under high uncertainty is challenging. We propose Marginal Advantage Estimation (MAE), a representation-level contribution estimation method for cooperative {MARL} under this structural isolation. MAE employs a synchronised feature-masking mechanism to evaluate marginal contributions without action-level counterfactual perturbations, thereby reducing variance and providing more informative learning signals. We provide a theoretical analysis of its bias–variance properties and demonstrate consistent performance improvements across 18 tasks in Partitioned MPE and Basilisk benchmarks over strong cooperative {MARL} baselines.
APA
Qiao, J., Cao, Z., Hu, S., Fan, M., Pratama, M. & Kowalczyk, R.. (2026). Multi-Agent RL with Invisible Collaborators: Marginal Advantage Estimation for Indirect Cooperation. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:5573-5600 Available from https://proceedings.mlr.press/v337/qiao26a.html.

Related Material