Turning Feature Attributions into Sufficient Explanations Using Conformal Prediction

Amr Alkhatib, Stephanie Lowry
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:382-401, 2026.

Abstract

Feature attribution methods are widely used for explaining image-based predictions, as they provide feature-level insights that can be intuitively visualized. However, such explanations often vary in their robustness and may fail to reflect the reasoning of the underlying black-box model faithfully. To address these limitations, we propose a novel conformal prediction-based approach that enables users to assign confidence levels to the important explanation regions. The proposed approach employs any feature-attribution method to identify a subset of salient features sufficient to preserve the model’s prediction, regardless of the information carried by the excluded features, without demanding access to ground-truth explanations for calibration. Four conformity functions are proposed to quantify the extent to which explanations conform to the model’s predictions. The approach is empirically evaluated using five explainers across six image datasets. The empirical results demonstrate that FastSHAP consistently outperforms the competing methods in terms of both fidelity and informational efficiency, the latter measured by the size of the explanation regions. Furthermore, the results reveal that conformity measures based on super-pixels are more effective than their pixel-wise counterparts.

Cite this Paper


BibTeX
@InProceedings{pmlr-v329-alkhatib26a, title = {Turning Feature Attributions into Sufficient Explanations Using Conformal Prediction}, author = {Alkhatib, Amr and Lowry, Stephanie}, booktitle = {Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications}, pages = {382--401}, year = {2026}, editor = {Ahlberg, Ernst and Johansson, Ulf and Boström, Henrik and Carlevaro, Alberto and Hallberg Szabadváry, Johan and Carlsson, Lars}, volume = {329}, series = {Proceedings of Machine Learning Research}, month = {02--04 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v329/main/assets/alkhatib26a/alkhatib26a.pdf}, url = {https://proceedings.mlr.press/v329/alkhatib26a.html}, abstract = {Feature attribution methods are widely used for explaining image-based predictions, as they provide feature-level insights that can be intuitively visualized. However, such explanations often vary in their robustness and may fail to reflect the reasoning of the underlying black-box model faithfully. To address these limitations, we propose a novel conformal prediction-based approach that enables users to assign confidence levels to the important explanation regions. The proposed approach employs any feature-attribution method to identify a subset of salient features sufficient to preserve the model’s prediction, regardless of the information carried by the excluded features, without demanding access to ground-truth explanations for calibration. Four conformity functions are proposed to quantify the extent to which explanations conform to the model’s predictions. The approach is empirically evaluated using five explainers across six image datasets. The empirical results demonstrate that FastSHAP consistently outperforms the competing methods in terms of both fidelity and informational efficiency, the latter measured by the size of the explanation regions. Furthermore, the results reveal that conformity measures based on super-pixels are more effective than their pixel-wise counterparts.} }
Endnote
%0 Conference Paper %T Turning Feature Attributions into Sufficient Explanations Using Conformal Prediction %A Amr Alkhatib %A Stephanie Lowry %B Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications %C Proceedings of Machine Learning Research %D 2026 %E Ernst Ahlberg %E Ulf Johansson %E Henrik Boström %E Alberto Carlevaro %E Johan Hallberg Szabadváry %E Lars Carlsson %F pmlr-v329-alkhatib26a %I PMLR %P 382--401 %U https://proceedings.mlr.press/v329/alkhatib26a.html %V 329 %X Feature attribution methods are widely used for explaining image-based predictions, as they provide feature-level insights that can be intuitively visualized. However, such explanations often vary in their robustness and may fail to reflect the reasoning of the underlying black-box model faithfully. To address these limitations, we propose a novel conformal prediction-based approach that enables users to assign confidence levels to the important explanation regions. The proposed approach employs any feature-attribution method to identify a subset of salient features sufficient to preserve the model’s prediction, regardless of the information carried by the excluded features, without demanding access to ground-truth explanations for calibration. Four conformity functions are proposed to quantify the extent to which explanations conform to the model’s predictions. The approach is empirically evaluated using five explainers across six image datasets. The empirical results demonstrate that FastSHAP consistently outperforms the competing methods in terms of both fidelity and informational efficiency, the latter measured by the size of the explanation regions. Furthermore, the results reveal that conformity measures based on super-pixels are more effective than their pixel-wise counterparts.
APA
Alkhatib, A. & Lowry, S.. (2026). Turning Feature Attributions into Sufficient Explanations Using Conformal Prediction. Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, in Proceedings of Machine Learning Research 329:382-401 Available from https://proceedings.mlr.press/v329/alkhatib26a.html.

Related Material