The Price of Valid Inference After Causal Discovery

Dongxin Guo, Jikun Wu, Siu Ming Yiu
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1804-1822, 2026.

Abstract

Causal effects estimated by covariate adjustment on a discovered graph have invalid confidence intervals, because the graph-selection step is ignored. We develop a unified selective-inference framework. For constraint-based discovery with latent confounders ({FCI}) under Gaussianity, we prove that the selection event, given the execution trace and signs, is a polyhedron in {Fisher}-z space; along the inference direction it becomes polynomial/rational constraints solving to a union of intervals. Inverting the truncated-{Gaussian} approximate pivot gives exact finite-sample $1-\alpha$ coverage, with known $\sigma^2$, of the adjustment functional $\gamma_1(S^\star)$ when $\{i\}\cup S^\star$ contains the outcome’s {Markov} blanket, and approximate coverage otherwise, with a non-vanishing distortion governed by the variance-ratio excess $\rho^2=\sigma^2_{j|X}/\sigma^2_{j|-j}-1$ (small under sparsity); the $\hat\sigma^2$ plug-in does not remove it. This is the first truncation-set characterization handling latent confounders. Coverage of the structural effect $\beta_{i\to j}$ is asymptotically valid up to the same distortion when the adjustment set is valid (exact when $\rho^2=0$), via high-dimensional {FCI} consistency under $d=o(\sqrt n)$. For unmodified GES, heuristic {Taylor} linearization indicates approximate coverage $1-\alpha-O(d^2/\sqrt n)$ when $n\gg d^4(\log d)^2$. In the nonparametric setting, $\Theta(n^{3/4})$-split discovery incurs width ratio $1+\Theta(n^{-1/4})$ versus a known-graph oracle.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-guo26a, title = {The Price of Valid Inference After Causal Discovery}, author = {Guo, Dongxin and Wu, Jikun and Yiu, Siu Ming}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {1804--1822}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/guo26a/guo26a.pdf}, url = {https://proceedings.mlr.press/v337/guo26a.html}, abstract = {Causal effects estimated by covariate adjustment on a discovered graph have invalid confidence intervals, because the graph-selection step is ignored. We develop a unified selective-inference framework. For constraint-based discovery with latent confounders ({FCI}) under Gaussianity, we prove that the selection event, given the execution trace and signs, is a polyhedron in {Fisher}-z space; along the inference direction it becomes polynomial/rational constraints solving to a union of intervals. Inverting the truncated-{Gaussian} approximate pivot gives exact finite-sample $1-\alpha$ coverage, with known $\sigma^2$, of the adjustment functional $\gamma_1(S^\star)$ when $\{i\}\cup S^\star$ contains the outcome’s {Markov} blanket, and approximate coverage otherwise, with a non-vanishing distortion governed by the variance-ratio excess $\rho^2=\sigma^2_{j|X}/\sigma^2_{j|-j}-1$ (small under sparsity); the $\hat\sigma^2$ plug-in does not remove it. This is the first truncation-set characterization handling latent confounders. Coverage of the structural effect $\beta_{i\to j}$ is asymptotically valid up to the same distortion when the adjustment set is valid (exact when $\rho^2=0$), via high-dimensional {FCI} consistency under $d=o(\sqrt n)$. For unmodified GES, heuristic {Taylor} linearization indicates approximate coverage $1-\alpha-O(d^2/\sqrt n)$ when $n\gg d^4(\log d)^2$. In the nonparametric setting, $\Theta(n^{3/4})$-split discovery incurs width ratio $1+\Theta(n^{-1/4})$ versus a known-graph oracle.} }
Endnote
%0 Conference Paper %T The Price of Valid Inference After Causal Discovery %A Dongxin Guo %A Jikun Wu %A Siu Ming Yiu %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-guo26a %I PMLR %P 1804--1822 %U https://proceedings.mlr.press/v337/guo26a.html %V 337 %X Causal effects estimated by covariate adjustment on a discovered graph have invalid confidence intervals, because the graph-selection step is ignored. We develop a unified selective-inference framework. For constraint-based discovery with latent confounders ({FCI}) under Gaussianity, we prove that the selection event, given the execution trace and signs, is a polyhedron in {Fisher}-z space; along the inference direction it becomes polynomial/rational constraints solving to a union of intervals. Inverting the truncated-{Gaussian} approximate pivot gives exact finite-sample $1-\alpha$ coverage, with known $\sigma^2$, of the adjustment functional $\gamma_1(S^\star)$ when $\{i\}\cup S^\star$ contains the outcome’s {Markov} blanket, and approximate coverage otherwise, with a non-vanishing distortion governed by the variance-ratio excess $\rho^2=\sigma^2_{j|X}/\sigma^2_{j|-j}-1$ (small under sparsity); the $\hat\sigma^2$ plug-in does not remove it. This is the first truncation-set characterization handling latent confounders. Coverage of the structural effect $\beta_{i\to j}$ is asymptotically valid up to the same distortion when the adjustment set is valid (exact when $\rho^2=0$), via high-dimensional {FCI} consistency under $d=o(\sqrt n)$. For unmodified GES, heuristic {Taylor} linearization indicates approximate coverage $1-\alpha-O(d^2/\sqrt n)$ when $n\gg d^4(\log d)^2$. In the nonparametric setting, $\Theta(n^{3/4})$-split discovery incurs width ratio $1+\Theta(n^{-1/4})$ versus a known-graph oracle.
APA
Guo, D., Wu, J. & Yiu, S.M.. (2026). The Price of Valid Inference After Causal Discovery. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:1804-1822 Available from https://proceedings.mlr.press/v337/guo26a.html.

Related Material