IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery

Ivaxi Sheth, Zhijing Jin, Bryan Wilder, Dominik Janzing, Mario Fritz
Proceedings of the Fifth Conference on Causal Learning and Reasoning, PMLR 323:1293-1317, 2026.

Abstract

In the presence of confounding between an endogenous variable and the outcome, instrumental variables (IVs) are used to isolate the causal effect of the endogenous variable. Identifying valid instruments requires interdisciplinary knowledge, creativity, and contextual understanding, making it a non-trivial task. In this paper, we investigate whether large language models (LLMs) can aid in this task. We perform a two-stage evaluation framework. First, we test whether LLMs can recover well-established instruments from the literature, assessing their ability to replicate standard reasoning. Second, we evaluate whether LLMs can identify and avoid instruments that have been empirically or theoretically discredited. Building on these results, we introduce IV Co-Scientist, a multi-agent system that proposes, critiques, and refines IVs for a given treatment–outcome pair. We also introduce a statistical test to contextualize consistency in the absence of ground truth. Our results show the potential of LLMs to discover novel valid instrumental variables from a large observational database.

Cite this Paper


BibTeX
@InProceedings{pmlr-v323-sheth26a, title = {IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery}, author = {Sheth, Ivaxi and Jin, Zhijing and Wilder, Bryan and Janzing, Dominik and Fritz, Mario}, booktitle = {Proceedings of the Fifth Conference on Causal Learning and Reasoning}, pages = {1293--1317}, year = {2026}, editor = {Mazaheri, Bijan and Hanson, Niels Richard}, volume = {323}, series = {Proceedings of Machine Learning Research}, month = {06--08 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v323/main/assets/sheth26a/sheth26a.pdf}, url = {https://proceedings.mlr.press/v323/sheth26a.html}, abstract = {In the presence of confounding between an endogenous variable and the outcome, instrumental variables (IVs) are used to isolate the causal effect of the endogenous variable. Identifying valid instruments requires interdisciplinary knowledge, creativity, and contextual understanding, making it a non-trivial task. In this paper, we investigate whether large language models (LLMs) can aid in this task. We perform a two-stage evaluation framework. First, we test whether LLMs can recover well-established instruments from the literature, assessing their ability to replicate standard reasoning. Second, we evaluate whether LLMs can identify and avoid instruments that have been empirically or theoretically discredited. Building on these results, we introduce IV Co-Scientist, a multi-agent system that proposes, critiques, and refines IVs for a given treatment–outcome pair. We also introduce a statistical test to contextualize consistency in the absence of ground truth. Our results show the potential of LLMs to discover novel valid instrumental variables from a large observational database.} }
Endnote
%0 Conference Paper %T IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery %A Ivaxi Sheth %A Zhijing Jin %A Bryan Wilder %A Dominik Janzing %A Mario Fritz %B Proceedings of the Fifth Conference on Causal Learning and Reasoning %C Proceedings of Machine Learning Research %D 2026 %E Bijan Mazaheri %E Niels Richard Hanson %F pmlr-v323-sheth26a %I PMLR %P 1293--1317 %U https://proceedings.mlr.press/v323/sheth26a.html %V 323 %X In the presence of confounding between an endogenous variable and the outcome, instrumental variables (IVs) are used to isolate the causal effect of the endogenous variable. Identifying valid instruments requires interdisciplinary knowledge, creativity, and contextual understanding, making it a non-trivial task. In this paper, we investigate whether large language models (LLMs) can aid in this task. We perform a two-stage evaluation framework. First, we test whether LLMs can recover well-established instruments from the literature, assessing their ability to replicate standard reasoning. Second, we evaluate whether LLMs can identify and avoid instruments that have been empirically or theoretically discredited. Building on these results, we introduce IV Co-Scientist, a multi-agent system that proposes, critiques, and refines IVs for a given treatment–outcome pair. We also introduce a statistical test to contextualize consistency in the absence of ground truth. Our results show the potential of LLMs to discover novel valid instrumental variables from a large observational database.
APA
Sheth, I., Jin, Z., Wilder, B., Janzing, D. & Fritz, M.. (2026). IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery. Proceedings of the Fifth Conference on Causal Learning and Reasoning, in Proceedings of Machine Learning Research 323:1293-1317 Available from https://proceedings.mlr.press/v323/sheth26a.html.

Related Material