Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach

Zihao Li, Fabrizio Russo
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3631-3667, 2026.

Abstract

Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of interventions. While expert knowledge is required to construct principled causal graphs, many statistical methods have been proposed to leverage observational data with varying formal guarantees. Causal Assumption-based Argumentation (ABA) is a framework that uses symbolic reasoning to ensure correspondence between input constraints and output graphs, while offering a principled way to combine data and expertise. We explore the use of large language models ({LLMs}) as imperfect experts, eliciting semantic structural constraints from variable names and descriptions and integrating them with statistical evidence through Causal ABA. We propose ABAPC-{LLM} as a principled and conservative approach to hybrid data- and {LLM}-driven causal discovery. Additionally, we introduce an evaluation protocol to mitigate memorisation bias when assessing {LLMs} for causal discovery and show competitive performance on novel semantically grounded random benchmarks, as well as on standard small- and medium-sized benchmarks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-li26e, title = {Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach}, author = {Li, Zihao and Russo, Fabrizio}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3631--3667}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/li26e/li26e.pdf}, url = {https://proceedings.mlr.press/v337/li26e.html}, abstract = {Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of interventions. While expert knowledge is required to construct principled causal graphs, many statistical methods have been proposed to leverage observational data with varying formal guarantees. Causal Assumption-based Argumentation (ABA) is a framework that uses symbolic reasoning to ensure correspondence between input constraints and output graphs, while offering a principled way to combine data and expertise. We explore the use of large language models ({LLMs}) as imperfect experts, eliciting semantic structural constraints from variable names and descriptions and integrating them with statistical evidence through Causal ABA. We propose ABAPC-{LLM} as a principled and conservative approach to hybrid data- and {LLM}-driven causal discovery. Additionally, we introduce an evaluation protocol to mitigate memorisation bias when assessing {LLMs} for causal discovery and show competitive performance on novel semantically grounded random benchmarks, as well as on standard small- and medium-sized benchmarks.} }
Endnote
%0 Conference Paper %T Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach %A Zihao Li %A Fabrizio Russo %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-li26e %I PMLR %P 3631--3667 %U https://proceedings.mlr.press/v337/li26e.html %V 337 %X Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of interventions. While expert knowledge is required to construct principled causal graphs, many statistical methods have been proposed to leverage observational data with varying formal guarantees. Causal Assumption-based Argumentation (ABA) is a framework that uses symbolic reasoning to ensure correspondence between input constraints and output graphs, while offering a principled way to combine data and expertise. We explore the use of large language models ({LLMs}) as imperfect experts, eliciting semantic structural constraints from variable names and descriptions and integrating them with statistical evidence through Causal ABA. We propose ABAPC-{LLM} as a principled and conservative approach to hybrid data- and {LLM}-driven causal discovery. Additionally, we introduce an evaluation protocol to mitigate memorisation bias when assessing {LLMs} for causal discovery and show competitive performance on novel semantically grounded random benchmarks, as well as on standard small- and medium-sized benchmarks.
APA
Li, Z. & Russo, F.. (2026). Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3631-3667 Available from https://proceedings.mlr.press/v337/li26e.html.

Related Material