The relative value of interventional and observational samples in Bayesian Causal Linear Gaussian Models

Valentinian Mihai Lungu, Anish Dhir, Mark van der Wilk, Ioannis Kontoyiannis
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:4067-4099, 2026.

Abstract

We investigate the asymptotic properties of {Bayesian} bivariate causal discovery for {Gaussian} Linear Structural Equation Models ({SEMs}) with heteroscedastic noise. We demonstrate that with purely observational data, the posterior distribution over the models fails to consistently identify the true causal structure—a consequence of the fundamental non-identifiability within the {Markov} Equivalence Class. Specifically, if the true generating mechanism corresponds to a connected graph ($A \rightarrow B$ or $B \rightarrow A$), the asymptotic behavior of the posterior is given by the ratio between the prior on the true model and the push-forward prior of the alternative. In contrast, for the independence model, we establish that the posterior concentrates at a stochastic polynomial rate of $O_p(n^{-1/2})$. To resolve this non-identifiability, we incorporate $m$ interventional samples and characterize the concentration rates as a function of the observational-to-total sample ratio, $\eta$. We identify a sharp \textit{concentration dichotomy}: while the independence graph maintains a polynomial $O_p(N^{-1/2})$ rate (where $N = n+m$), connected graphs undergo a phase transition to exponentially fast convergence. This highlights an \textit{exponential} relative importance between the two data types, as altering the amount of one data type directly changes the exponent governing the concentration speed. We derive explicit formulae for the exponential decay rates and provide precise conditions under which mixing observational and interventional data optimizes concentration speed. Finally, our theoretical findings are validated through empirical simulations in {Bayesian} {Gaussian} equivalent (BGe)-style prior specifications offering a principled foundation for experimental design in {Bayesian} causal discovery.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-lungu26a, title = {The relative value of interventional and observational samples in {Bayesian} Causal Linear {Gaussian} Models}, author = {Lungu, Valentinian Mihai and Dhir, Anish and van der Wilk, Mark and Kontoyiannis, Ioannis}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {4067--4099}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/lungu26a/lungu26a.pdf}, url = {https://proceedings.mlr.press/v337/lungu26a.html}, abstract = {We investigate the asymptotic properties of {Bayesian} bivariate causal discovery for {Gaussian} Linear Structural Equation Models ({SEMs}) with heteroscedastic noise. We demonstrate that with purely observational data, the posterior distribution over the models fails to consistently identify the true causal structure—a consequence of the fundamental non-identifiability within the {Markov} Equivalence Class. Specifically, if the true generating mechanism corresponds to a connected graph ($A \rightarrow B$ or $B \rightarrow A$), the asymptotic behavior of the posterior is given by the ratio between the prior on the true model and the push-forward prior of the alternative. In contrast, for the independence model, we establish that the posterior concentrates at a stochastic polynomial rate of $O_p(n^{-1/2})$. To resolve this non-identifiability, we incorporate $m$ interventional samples and characterize the concentration rates as a function of the observational-to-total sample ratio, $\eta$. We identify a sharp \textit{concentration dichotomy}: while the independence graph maintains a polynomial $O_p(N^{-1/2})$ rate (where $N = n+m$), connected graphs undergo a phase transition to exponentially fast convergence. This highlights an \textit{exponential} relative importance between the two data types, as altering the amount of one data type directly changes the exponent governing the concentration speed. We derive explicit formulae for the exponential decay rates and provide precise conditions under which mixing observational and interventional data optimizes concentration speed. Finally, our theoretical findings are validated through empirical simulations in {Bayesian} {Gaussian} equivalent (BGe)-style prior specifications offering a principled foundation for experimental design in {Bayesian} causal discovery.} }
Endnote
%0 Conference Paper %T The relative value of interventional and observational samples in Bayesian Causal Linear Gaussian Models %A Valentinian Mihai Lungu %A Anish Dhir %A Mark van der Wilk %A Ioannis Kontoyiannis %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-lungu26a %I PMLR %P 4067--4099 %U https://proceedings.mlr.press/v337/lungu26a.html %V 337 %X We investigate the asymptotic properties of {Bayesian} bivariate causal discovery for {Gaussian} Linear Structural Equation Models ({SEMs}) with heteroscedastic noise. We demonstrate that with purely observational data, the posterior distribution over the models fails to consistently identify the true causal structure—a consequence of the fundamental non-identifiability within the {Markov} Equivalence Class. Specifically, if the true generating mechanism corresponds to a connected graph ($A \rightarrow B$ or $B \rightarrow A$), the asymptotic behavior of the posterior is given by the ratio between the prior on the true model and the push-forward prior of the alternative. In contrast, for the independence model, we establish that the posterior concentrates at a stochastic polynomial rate of $O_p(n^{-1/2})$. To resolve this non-identifiability, we incorporate $m$ interventional samples and characterize the concentration rates as a function of the observational-to-total sample ratio, $\eta$. We identify a sharp \textit{concentration dichotomy}: while the independence graph maintains a polynomial $O_p(N^{-1/2})$ rate (where $N = n+m$), connected graphs undergo a phase transition to exponentially fast convergence. This highlights an \textit{exponential} relative importance between the two data types, as altering the amount of one data type directly changes the exponent governing the concentration speed. We derive explicit formulae for the exponential decay rates and provide precise conditions under which mixing observational and interventional data optimizes concentration speed. Finally, our theoretical findings are validated through empirical simulations in {Bayesian} {Gaussian} equivalent (BGe)-style prior specifications offering a principled foundation for experimental design in {Bayesian} causal discovery.
APA
Lungu, V.M., Dhir, A., van der Wilk, M. & Kontoyiannis, I.. (2026). The relative value of interventional and observational samples in Bayesian Causal Linear Gaussian Models. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:4067-4099 Available from https://proceedings.mlr.press/v337/lungu26a.html.

Related Material