Causal Foundations of Collective Agency

Frederik Hytting Jørgensen, Sebastian Weichwald, Lewis Hammond
Proceedings of the Fifth Conference on Causal Learning and Reasoning, PMLR 323:1140-1170, 2026.

Abstract

A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining when a group of agents can be viewed as a unified collective agent is a foundational question in the study of interactions and incentives in both biological and artificial systems. We adopt a behavioral perspective in answering this question, ascribing collective agency to a group when viewing the group’s joint actions as rational and goal-directed successfully predicts its behavior. We formalize this perspective on collective agency using causal games (Hammond et al., 2023) – which are causal models of strategic, multi-agent interactions – and causal abstraction (Rubenstein et al., 2017; Beckers and Halpern, 2019) – which formalizes when a simple, high-level model faithfully captures a more complex, low-level model. We use this framework to solve a puzzle regarding multi-agent incentives in actor-critic models and to make quantitative assessments of the degree of collective agency exhibited by different voting mechanisms. Our framework aims to provide a foundation for theoretical and empirical work to understand, predict, and control emergent collective agents in multi-agent AI systems

Cite this Paper


BibTeX
@InProceedings{pmlr-v323-jorgensen26a, title = {Causal Foundations of Collective Agency}, author = {J{\o}rgensen, Frederik Hytting and Weichwald, Sebastian and Hammond, Lewis}, booktitle = {Proceedings of the Fifth Conference on Causal Learning and Reasoning}, pages = {1140--1170}, year = {2026}, editor = {Mazaheri, Bijan and Hanson, Niels Richard}, volume = {323}, series = {Proceedings of Machine Learning Research}, month = {06--08 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v323/main/assets/jorgensen26a/jorgensen26a.pdf}, url = {https://proceedings.mlr.press/v323/jorgensen26a.html}, abstract = {A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining when a group of agents can be viewed as a unified collective agent is a foundational question in the study of interactions and incentives in both biological and artificial systems. We adopt a behavioral perspective in answering this question, ascribing collective agency to a group when viewing the group’s joint actions as rational and goal-directed successfully predicts its behavior. We formalize this perspective on collective agency using causal games (Hammond et al., 2023) – which are causal models of strategic, multi-agent interactions – and causal abstraction (Rubenstein et al., 2017; Beckers and Halpern, 2019) – which formalizes when a simple, high-level model faithfully captures a more complex, low-level model. We use this framework to solve a puzzle regarding multi-agent incentives in actor-critic models and to make quantitative assessments of the degree of collective agency exhibited by different voting mechanisms. Our framework aims to provide a foundation for theoretical and empirical work to understand, predict, and control emergent collective agents in multi-agent AI systems} }
Endnote
%0 Conference Paper %T Causal Foundations of Collective Agency %A Frederik Hytting Jørgensen %A Sebastian Weichwald %A Lewis Hammond %B Proceedings of the Fifth Conference on Causal Learning and Reasoning %C Proceedings of Machine Learning Research %D 2026 %E Bijan Mazaheri %E Niels Richard Hanson %F pmlr-v323-jorgensen26a %I PMLR %P 1140--1170 %U https://proceedings.mlr.press/v323/jorgensen26a.html %V 323 %X A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining when a group of agents can be viewed as a unified collective agent is a foundational question in the study of interactions and incentives in both biological and artificial systems. We adopt a behavioral perspective in answering this question, ascribing collective agency to a group when viewing the group’s joint actions as rational and goal-directed successfully predicts its behavior. We formalize this perspective on collective agency using causal games (Hammond et al., 2023) – which are causal models of strategic, multi-agent interactions – and causal abstraction (Rubenstein et al., 2017; Beckers and Halpern, 2019) – which formalizes when a simple, high-level model faithfully captures a more complex, low-level model. We use this framework to solve a puzzle regarding multi-agent incentives in actor-critic models and to make quantitative assessments of the degree of collective agency exhibited by different voting mechanisms. Our framework aims to provide a foundation for theoretical and empirical work to understand, predict, and control emergent collective agents in multi-agent AI systems
APA
Jørgensen, F.H., Weichwald, S. & Hammond, L.. (2026). Causal Foundations of Collective Agency. Proceedings of the Fifth Conference on Causal Learning and Reasoning, in Proceedings of Machine Learning Research 323:1140-1170 Available from https://proceedings.mlr.press/v323/jorgensen26a.html.

Related Material