Unveiling Causal Calibration: How LLM Scale Influences Statistical Information Interpolation

Markus Englberger, Devendra Singh Dhami
Proceedings of the Fifth Conference on Causal Learning and Reasoning, PMLR 323:757-775, 2026.

Abstract

There are two ways an LLM can learn statistical and causal knowledge during training: either by explicit statements about the causal scenario in the training corpus (e.g., “the probability of developing lung cancer is 20% if one has smoked for at least 10 years”, “X and Y are independent conditioned on Z”, “X is a cause of Y” etc.) or by implicit presentations of the same information via tabular data from surveys or scenario descriptions. In this paper, we probe whether the LLM internally unifies these two sources of information. To this end, we test whether fine-tuning LLMs on implicit presentations of statistical knowledge influences the LLM’s answers when explicitly prompted for such knowledge. By creating in-context learning tasks where both explicit and implicit statistical information are present, we also test whether language models can unify these two modes during inference. Experiments suggest that larger language models have some shared method of storing these two sources of information, whereas smaller language models seem to be less capable of unifying these sources.

Cite this Paper


BibTeX
@InProceedings{pmlr-v323-englberger26a, title = {Unveiling Causal Calibration: How LLM Scale Influences Statistical Information Interpolation}, author = {Englberger, Markus and Dhami, Devendra Singh}, booktitle = {Proceedings of the Fifth Conference on Causal Learning and Reasoning}, pages = {757--775}, year = {2026}, editor = {Mazaheri, Bijan and Hanson, Niels Richard}, volume = {323}, series = {Proceedings of Machine Learning Research}, month = {06--08 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v323/main/assets/englberger26a/englberger26a.pdf}, url = {https://proceedings.mlr.press/v323/englberger26a.html}, abstract = {There are two ways an LLM can learn statistical and causal knowledge during training: either by explicit statements about the causal scenario in the training corpus (e.g., “the probability of developing lung cancer is 20% if one has smoked for at least 10 years”, “X and Y are independent conditioned on Z”, “X is a cause of Y” etc.) or by implicit presentations of the same information via tabular data from surveys or scenario descriptions. In this paper, we probe whether the LLM internally unifies these two sources of information. To this end, we test whether fine-tuning LLMs on implicit presentations of statistical knowledge influences the LLM’s answers when explicitly prompted for such knowledge. By creating in-context learning tasks where both explicit and implicit statistical information are present, we also test whether language models can unify these two modes during inference. Experiments suggest that larger language models have some shared method of storing these two sources of information, whereas smaller language models seem to be less capable of unifying these sources.} }
Endnote
%0 Conference Paper %T Unveiling Causal Calibration: How LLM Scale Influences Statistical Information Interpolation %A Markus Englberger %A Devendra Singh Dhami %B Proceedings of the Fifth Conference on Causal Learning and Reasoning %C Proceedings of Machine Learning Research %D 2026 %E Bijan Mazaheri %E Niels Richard Hanson %F pmlr-v323-englberger26a %I PMLR %P 757--775 %U https://proceedings.mlr.press/v323/englberger26a.html %V 323 %X There are two ways an LLM can learn statistical and causal knowledge during training: either by explicit statements about the causal scenario in the training corpus (e.g., “the probability of developing lung cancer is 20% if one has smoked for at least 10 years”, “X and Y are independent conditioned on Z”, “X is a cause of Y” etc.) or by implicit presentations of the same information via tabular data from surveys or scenario descriptions. In this paper, we probe whether the LLM internally unifies these two sources of information. To this end, we test whether fine-tuning LLMs on implicit presentations of statistical knowledge influences the LLM’s answers when explicitly prompted for such knowledge. By creating in-context learning tasks where both explicit and implicit statistical information are present, we also test whether language models can unify these two modes during inference. Experiments suggest that larger language models have some shared method of storing these two sources of information, whereas smaller language models seem to be less capable of unifying these sources.
APA
Englberger, M. & Dhami, D.S.. (2026). Unveiling Causal Calibration: How LLM Scale Influences Statistical Information Interpolation. Proceedings of the Fifth Conference on Causal Learning and Reasoning, in Proceedings of Machine Learning Research 323:757-775 Available from https://proceedings.mlr.press/v323/englberger26a.html.

Related Material