Conformal Selection of Counterfactual Explanations: From Minimal to Robust Recourse

Ulf Johansson, Aicha Maalej, Niclas Ståhl, Henrik Boström
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:529-547, 2026.

Abstract

Counterfactual explanations are commonly generated by identifying minimally perturbed instances that change a model prediction. However, minimal perturbations do not necessarily correspond to realistic, representative, or robust forms of recourse. In particular, counterfactuals located close to the decision boundary may satisfy the desired prediction outcome while still exhibiting atypical feature configurations or relying on unusual model reasoning patterns. This paper proposes a conformal framework for selecting counterfactual explanations based on their conformity to the target-class distribution. Rather than selecting the closest generated counterfactual, the proposed approach selects the candidate with highest conformal p-value according to conformal anomaly detection. Conformity is evaluated both in feature space and in SHAP space, enabling robustness to be assessed not only in terms of geometric similarity, but also in terms of explanatory consistency. Experiments on twelve benchmark datasets demonstrate that substantially more conforming counterfactual explanations can often be obtained with only modest increases in perturbation magnitude. In particular, the results show that counterfactuals with similar feature-space proximity may nevertheless differ substantially in SHAP-space conformity, suggesting that minimally perturbed counterfactuals may rely on atypical explanatory patterns. Qualitative examples further illustrate how conformity-based selection can produce explanations that appear more representative of genuine target-class instances. The proposed framework further highlights how conformity-based reasoning can extend beyond predictive uncertainty estimation and provide principled tools for assessing the quality and consistency of AI explanations.

Cite this Paper


BibTeX
@InProceedings{pmlr-v329-johansson26a, title = {Conformal Selection of Counterfactual Explanations: From Minimal to Robust Recourse}, author = {Johansson, Ulf and Maalej, Aicha and St{\aa}hl, Niclas and Bostr{\"o}m, Henrik}, booktitle = {Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications}, pages = {529--547}, year = {2026}, editor = {Ahlberg, Ernst and Johansson, Ulf and Boström, Henrik and Carlevaro, Alberto and Hallberg Szabadváry, Johan and Carlsson, Lars}, volume = {329}, series = {Proceedings of Machine Learning Research}, month = {02--04 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v329/main/assets/johansson26a/johansson26a.pdf}, url = {https://proceedings.mlr.press/v329/johansson26a.html}, abstract = {Counterfactual explanations are commonly generated by identifying minimally perturbed instances that change a model prediction. However, minimal perturbations do not necessarily correspond to realistic, representative, or robust forms of recourse. In particular, counterfactuals located close to the decision boundary may satisfy the desired prediction outcome while still exhibiting atypical feature configurations or relying on unusual model reasoning patterns. This paper proposes a conformal framework for selecting counterfactual explanations based on their conformity to the target-class distribution. Rather than selecting the closest generated counterfactual, the proposed approach selects the candidate with highest conformal p-value according to conformal anomaly detection. Conformity is evaluated both in feature space and in SHAP space, enabling robustness to be assessed not only in terms of geometric similarity, but also in terms of explanatory consistency. Experiments on twelve benchmark datasets demonstrate that substantially more conforming counterfactual explanations can often be obtained with only modest increases in perturbation magnitude. In particular, the results show that counterfactuals with similar feature-space proximity may nevertheless differ substantially in SHAP-space conformity, suggesting that minimally perturbed counterfactuals may rely on atypical explanatory patterns. Qualitative examples further illustrate how conformity-based selection can produce explanations that appear more representative of genuine target-class instances. The proposed framework further highlights how conformity-based reasoning can extend beyond predictive uncertainty estimation and provide principled tools for assessing the quality and consistency of AI explanations.} }
Endnote
%0 Conference Paper %T Conformal Selection of Counterfactual Explanations: From Minimal to Robust Recourse %A Ulf Johansson %A Aicha Maalej %A Niclas Ståhl %A Henrik Boström %B Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications %C Proceedings of Machine Learning Research %D 2026 %E Ernst Ahlberg %E Ulf Johansson %E Henrik Boström %E Alberto Carlevaro %E Johan Hallberg Szabadváry %E Lars Carlsson %F pmlr-v329-johansson26a %I PMLR %P 529--547 %U https://proceedings.mlr.press/v329/johansson26a.html %V 329 %X Counterfactual explanations are commonly generated by identifying minimally perturbed instances that change a model prediction. However, minimal perturbations do not necessarily correspond to realistic, representative, or robust forms of recourse. In particular, counterfactuals located close to the decision boundary may satisfy the desired prediction outcome while still exhibiting atypical feature configurations or relying on unusual model reasoning patterns. This paper proposes a conformal framework for selecting counterfactual explanations based on their conformity to the target-class distribution. Rather than selecting the closest generated counterfactual, the proposed approach selects the candidate with highest conformal p-value according to conformal anomaly detection. Conformity is evaluated both in feature space and in SHAP space, enabling robustness to be assessed not only in terms of geometric similarity, but also in terms of explanatory consistency. Experiments on twelve benchmark datasets demonstrate that substantially more conforming counterfactual explanations can often be obtained with only modest increases in perturbation magnitude. In particular, the results show that counterfactuals with similar feature-space proximity may nevertheless differ substantially in SHAP-space conformity, suggesting that minimally perturbed counterfactuals may rely on atypical explanatory patterns. Qualitative examples further illustrate how conformity-based selection can produce explanations that appear more representative of genuine target-class instances. The proposed framework further highlights how conformity-based reasoning can extend beyond predictive uncertainty estimation and provide principled tools for assessing the quality and consistency of AI explanations.
APA
Johansson, U., Maalej, A., Ståhl, N. & Boström, H.. (2026). Conformal Selection of Counterfactual Explanations: From Minimal to Robust Recourse. Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, in Proceedings of Machine Learning Research 329:529-547 Available from https://proceedings.mlr.press/v329/johansson26a.html.

Related Material