CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?

Sawal Acharya, Terry Jingchen Zhang, Andrew Kim, Rahul Babu Shrestha, Xianlin Sun, Pepijn Cobben, Maximilian Mordig, Jacob T. Emmerson, Anahita Haghighat, Furkan Danisman, Yuen Chen, Clijo Jose, Andrei Ioan Muresanu, Justin Cui, Jiarui Liu, Yahang Qi, Punya Syon Pandey, Yinya Huang, Bernhard Schölkopf, Zhijing Jin
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:332-384, 2026.

Abstract

Identifying and estimating causal relationships from data is a crucial component of empirical research. While large language model-powered tools have shown potential for assisting research workflows, their ability to perform end-to-end causal inference remains underexplored. We introduce CauSciBench, a benchmark that puts LLM-powered tools to the test on causality- driven research questions. Unlike previous related benchmarks that focus on coding alone, CauSciBench enables evaluation across the full pipeline of causal inference: from method and variable selection to computation of causal effects and statistical interpretation in the context of real-world research problems. We evaluated 7 frontier models on over 300 queries derived from scientific publications, textbook problems, sem- inal datasets, and synthetic scenarios. Results show that models consistently perform worse on real datasets, with the key bottleneck being the selection of an appropriate causal inference method.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-acharya26a, title = {{C}au{S}ci{B}ench: Can {LLM}s Automate Causal Inference in Real-World Scientific Research?}, author = {Acharya, Sawal and Zhang, Terry Jingchen and Kim, Andrew and Shrestha, Rahul Babu and Sun, Xianlin and Cobben, Pepijn and Mordig, Maximilian and Emmerson, Jacob T. and Haghighat, Anahita and Danisman, Furkan and Chen, Yuen and Jose, Clijo and Muresanu, Andrei Ioan and Cui, Justin and Liu, Jiarui and Qi, Yahang and Pandey, Punya Syon and Huang, Yinya and Sch\"{o}lkopf, Bernhard and Jin, Zhijing}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {332--384}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/acharya26a/acharya26a.pdf}, url = {https://proceedings.mlr.press/v306/acharya26a.html}, abstract = {Identifying and estimating causal relationships from data is a crucial component of empirical research. While large language model-powered tools have shown potential for assisting research workflows, their ability to perform end-to-end causal inference remains underexplored. We introduce CauSciBench, a benchmark that puts LLM-powered tools to the test on causality- driven research questions. Unlike previous related benchmarks that focus on coding alone, CauSciBench enables evaluation across the full pipeline of causal inference: from method and variable selection to computation of causal effects and statistical interpretation in the context of real-world research problems. We evaluated 7 frontier models on over 300 queries derived from scientific publications, textbook problems, sem- inal datasets, and synthetic scenarios. Results show that models consistently perform worse on real datasets, with the key bottleneck being the selection of an appropriate causal inference method.} }
Endnote
%0 Conference Paper %T CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? %A Sawal Acharya %A Terry Jingchen Zhang %A Andrew Kim %A Rahul Babu Shrestha %A Xianlin Sun %A Pepijn Cobben %A Maximilian Mordig %A Jacob T. Emmerson %A Anahita Haghighat %A Furkan Danisman %A Yuen Chen %A Clijo Jose %A Andrei Ioan Muresanu %A Justin Cui %A Jiarui Liu %A Yahang Qi %A Punya Syon Pandey %A Yinya Huang %A Bernhard Schölkopf %A Zhijing Jin %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-acharya26a %I PMLR %P 332--384 %U https://proceedings.mlr.press/v306/acharya26a.html %V 306 %X Identifying and estimating causal relationships from data is a crucial component of empirical research. While large language model-powered tools have shown potential for assisting research workflows, their ability to perform end-to-end causal inference remains underexplored. We introduce CauSciBench, a benchmark that puts LLM-powered tools to the test on causality- driven research questions. Unlike previous related benchmarks that focus on coding alone, CauSciBench enables evaluation across the full pipeline of causal inference: from method and variable selection to computation of causal effects and statistical interpretation in the context of real-world research problems. We evaluated 7 frontier models on over 300 queries derived from scientific publications, textbook problems, sem- inal datasets, and synthetic scenarios. Results show that models consistently perform worse on real datasets, with the key bottleneck being the selection of an appropriate causal inference method.
APA
Acharya, S., Zhang, T.J., Kim, A., Shrestha, R.B., Sun, X., Cobben, P., Mordig, M., Emmerson, J.T., Haghighat, A., Danisman, F., Chen, Y., Jose, C., Muresanu, A.I., Cui, J., Liu, J., Qi, Y., Pandey, P.S., Huang, Y., Schölkopf, B. & Jin, Z.. (2026). CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:332-384 Available from https://proceedings.mlr.press/v306/acharya26a.html.

Related Material