cc-Shapley: Measuring Multivariate Feature Importance Needs Causal Context

Jörg Martin, Stefan Haufe
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:4323-4352, 2026.

Abstract

Explainable artificial intelligence promises to yield insights into relevant features, thereby enabling humans to examine and scrutinize machine learning models or even facilitating scientific discovery. Considering the widespread technique of {Shapley} values, we find that purely data-driven operationalization of multivariate feature importance is unsuitable for such purposes. Even for simple problems with two features, spurious associations due to collider bias and suppression arise from considering one feature only in the observational context of the other, which can lead to misinterpretations. Causal knowledge about the data-generating process is required to identify and correct such misleading feature attributions. We propose cc-{Shapley} (causal context {Shapley}), an interventional modification of conventional observational {Shapley} values leveraging knowledge of the data’s causal structure, thereby analyzing the relevance of a feature in the causal context of the remaining features. We show theoretically that this eradicates spurious association induced by collider bias. We compare the behavior of {Shapley} and cc-{Shapley} values on various, synthetic, and real-world datasets. We observe nullification or reversal of associations when moving from observational to cc-{Shapley}.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-martin26a, title = {cc-{Shapley}: Measuring Multivariate Feature Importance Needs Causal Context}, author = {Martin, J\"{o}rg and Haufe, Stefan}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {4323--4352}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/martin26a/martin26a.pdf}, url = {https://proceedings.mlr.press/v337/martin26a.html}, abstract = {Explainable artificial intelligence promises to yield insights into relevant features, thereby enabling humans to examine and scrutinize machine learning models or even facilitating scientific discovery. Considering the widespread technique of {Shapley} values, we find that purely data-driven operationalization of multivariate feature importance is unsuitable for such purposes. Even for simple problems with two features, spurious associations due to collider bias and suppression arise from considering one feature only in the observational context of the other, which can lead to misinterpretations. Causal knowledge about the data-generating process is required to identify and correct such misleading feature attributions. We propose cc-{Shapley} (causal context {Shapley}), an interventional modification of conventional observational {Shapley} values leveraging knowledge of the data’s causal structure, thereby analyzing the relevance of a feature in the causal context of the remaining features. We show theoretically that this eradicates spurious association induced by collider bias. We compare the behavior of {Shapley} and cc-{Shapley} values on various, synthetic, and real-world datasets. We observe nullification or reversal of associations when moving from observational to cc-{Shapley}.} }
Endnote
%0 Conference Paper %T cc-Shapley: Measuring Multivariate Feature Importance Needs Causal Context %A Jörg Martin %A Stefan Haufe %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-martin26a %I PMLR %P 4323--4352 %U https://proceedings.mlr.press/v337/martin26a.html %V 337 %X Explainable artificial intelligence promises to yield insights into relevant features, thereby enabling humans to examine and scrutinize machine learning models or even facilitating scientific discovery. Considering the widespread technique of {Shapley} values, we find that purely data-driven operationalization of multivariate feature importance is unsuitable for such purposes. Even for simple problems with two features, spurious associations due to collider bias and suppression arise from considering one feature only in the observational context of the other, which can lead to misinterpretations. Causal knowledge about the data-generating process is required to identify and correct such misleading feature attributions. We propose cc-{Shapley} (causal context {Shapley}), an interventional modification of conventional observational {Shapley} values leveraging knowledge of the data’s causal structure, thereby analyzing the relevance of a feature in the causal context of the remaining features. We show theoretically that this eradicates spurious association induced by collider bias. We compare the behavior of {Shapley} and cc-{Shapley} values on various, synthetic, and real-world datasets. We observe nullification or reversal of associations when moving from observational to cc-{Shapley}.
APA
Martin, J. & Haufe, S.. (2026). cc-Shapley: Measuring Multivariate Feature Importance Needs Causal Context. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:4323-4352 Available from https://proceedings.mlr.press/v337/martin26a.html.

Related Material