Missing Data as a Causal and Probabilistic Problem

Ilya Shpitser, Karthika Mohan UCLA, Judea Pearl UCLA
Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, PMLR R13:614-623, 2015.

Abstract

Causal inference is often phrased as a missing data problem – for every unit, only the response to observed treatment assignment is known, the response to other treatment assignments is not. In this paper, we extend the converse approach of (Mohan et al, 2013) of representing missing data problems to restricted causal models (where only interventions on missingness indicators are allowed). We further use this representation to leverage techniques developed for the problem of identification of causal effects to give a general criterion for cases where a joint distribution containing missing variables can be recovered from data actually observed, given assumptions on missingness mechanisms. This criterion is significantly more general than the commonly used “missing at random” (MAR) criterion, and generalizes past work which also exploits a graphical representation of missingness. In fact, the relationship of our criterion to MAR is not unlike the relationship between the ID algorithm for identification of causal effects (Tian and Pearl, 2002), (Shpitser and Pearl 2006), and conditional ignorability (Rosenbaum and Rubin, 1983).

Cite this Paper


BibTeX
@InProceedings{pmlr-vR13-shpitser15a, title = {Missing Data as a Causal and Probabilistic Problem}, author = {Shpitser, Ilya and UCLA, Karthika Mohan and UCLA, Judea Pearl}, booktitle = {Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence}, pages = {614--623}, year = {2015}, editor = {Meila, Marina and Heskes, Tom}, volume = {R13}, series = {Proceedings of Machine Learning Research}, month = {12--16 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r13/main/assets/shpitser15a/shpitser15a.pdf}, url = {https://proceedings.mlr.press/r13/shpitser15a.html}, abstract = {Causal inference is often phrased as a missing data problem – for every unit, only the response to observed treatment assignment is known, the response to other treatment assignments is not. In this paper, we extend the converse approach of (Mohan et al, 2013) of representing missing data problems to restricted causal models (where only interventions on missingness indicators are allowed). We further use this representation to leverage techniques developed for the problem of identification of causal effects to give a general criterion for cases where a joint distribution containing missing variables can be recovered from data actually observed, given assumptions on missingness mechanisms. This criterion is significantly more general than the commonly used “missing at random” (MAR) criterion, and generalizes past work which also exploits a graphical representation of missingness. In fact, the relationship of our criterion to MAR is not unlike the relationship between the ID algorithm for identification of causal effects (Tian and Pearl, 2002), (Shpitser and Pearl 2006), and conditional ignorability (Rosenbaum and Rubin, 1983).}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Missing Data as a Causal and Probabilistic Problem %A Ilya Shpitser %A Karthika Mohan UCLA %A Judea Pearl UCLA %B Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2015 %E Marina Meila %E Tom Heskes %F pmlr-vR13-shpitser15a %I PMLR %P 614--623 %U https://proceedings.mlr.press/r13/shpitser15a.html %V R13 %X Causal inference is often phrased as a missing data problem – for every unit, only the response to observed treatment assignment is known, the response to other treatment assignments is not. In this paper, we extend the converse approach of (Mohan et al, 2013) of representing missing data problems to restricted causal models (where only interventions on missingness indicators are allowed). We further use this representation to leverage techniques developed for the problem of identification of causal effects to give a general criterion for cases where a joint distribution containing missing variables can be recovered from data actually observed, given assumptions on missingness mechanisms. This criterion is significantly more general than the commonly used “missing at random” (MAR) criterion, and generalizes past work which also exploits a graphical representation of missingness. In fact, the relationship of our criterion to MAR is not unlike the relationship between the ID algorithm for identification of causal effects (Tian and Pearl, 2002), (Shpitser and Pearl 2006), and conditional ignorability (Rosenbaum and Rubin, 1983). %Z Reissued by PMLR on 04 October 2026.
APA
Shpitser, I., UCLA, K.M. & UCLA, J.P.. (2015). Missing Data as a Causal and Probabilistic Problem. Proceedings of the 31st Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R13:614-623 Available from https://proceedings.mlr.press/r13/shpitser15a.html. Reissued by PMLR on 04 October 2026.

Related Material