[edit]
Conformal Selection of Counterfactual Explanations: From Minimal to Robust Recourse
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:529-547, 2026.
Abstract
Counterfactual explanations are commonly generated by identifying minimally perturbed instances that change a model prediction. However, minimal perturbations do not necessarily correspond to realistic, representative, or robust forms of recourse. In particular, counterfactuals located close to the decision boundary may satisfy the desired prediction outcome while still exhibiting atypical feature configurations or relying on unusual model reasoning patterns. This paper proposes a conformal framework for selecting counterfactual explanations based on their conformity to the target-class distribution. Rather than selecting the closest generated counterfactual, the proposed approach selects the candidate with highest conformal p-value according to conformal anomaly detection. Conformity is evaluated both in feature space and in SHAP space, enabling robustness to be assessed not only in terms of geometric similarity, but also in terms of explanatory consistency. Experiments on twelve benchmark datasets demonstrate that substantially more conforming counterfactual explanations can often be obtained with only modest increases in perturbation magnitude. In particular, the results show that counterfactuals with similar feature-space proximity may nevertheless differ substantially in SHAP-space conformity, suggesting that minimally perturbed counterfactuals may rely on atypical explanatory patterns. Qualitative examples further illustrate how conformity-based selection can produce explanations that appear more representative of genuine target-class instances. The proposed framework further highlights how conformity-based reasoning can extend beyond predictive uncertainty estimation and provide principled tools for assessing the quality and consistency of AI explanations.