Guarded Explanations: Conformal-Style Filtering for Distribution-Aware Rule Conditions

Tuwe Löfström, Anders Hjort, Helena Löfström
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:422-438, 2026.

Abstract

In high-stakes decision support, post-hoc explanation methods are often used to estimate feature influence by perturbing inputs. However, a fundamental limitation of this approach is that candidate perturbations frequently fall outside the support of the calibration data, yielding unsupported and potentially misleading explanations. We introduce guarded explanations, an extension of the Calibrated Explanations framework. Guarded explanations apply a conformal-anomaly-detection-inspired filter to candidate perturbations, retaining a representative candidate only when its empirical guard score meets a user-chosen threshold $\varepsilon$. To facilitate this, we replace the standard binary discretiser with a multi-bin discretiser, generating interval-valued conditions alongside one-sided thresholds. Emitted perturbations are then scored using the standard Calibrated Explanations uncertainty-aware backend. In a synthetic constraint setting, we demonstrate that guarded factual explanations reduce the fraction of emitted rules whose representative values violate a known domain constraint from approximately 2.5% to near 0.0% at $\varepsilon$ = 0.2. Crucially, we isolate this effect from discretisation granularity alone by comparing against an unpruned multi-bin (no-guard) baseline. We frame guarded explanations as a representative-level plausibility filter, where the guard demonstrably reduces calibration-unsupported representative candidates in emitted rules. The guard score is a conformal-style empirical rank, not a standard conformal p-value, and $\varepsilon$ has no finite-sample error-rate guarantee for constructed perturbations.

Cite this Paper


BibTeX
@InProceedings{pmlr-v329-lofstrom26a, title = {Guarded Explanations: Conformal-Style Filtering for Distribution-Aware Rule Conditions}, author = {L{\"o}fstr{\"o}m, Tuwe and Hjort, Anders and L{\"o}fstr{\"o}m, Helena}, booktitle = {Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications}, pages = {422--438}, year = {2026}, editor = {Ahlberg, Ernst and Johansson, Ulf and Boström, Henrik and Carlevaro, Alberto and Hallberg Szabadváry, Johan and Carlsson, Lars}, volume = {329}, series = {Proceedings of Machine Learning Research}, month = {02--04 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v329/main/assets/lofstrom26a/lofstrom26a.pdf}, url = {https://proceedings.mlr.press/v329/lofstrom26a.html}, abstract = {In high-stakes decision support, post-hoc explanation methods are often used to estimate feature influence by perturbing inputs. However, a fundamental limitation of this approach is that candidate perturbations frequently fall outside the support of the calibration data, yielding unsupported and potentially misleading explanations. We introduce guarded explanations, an extension of the Calibrated Explanations framework. Guarded explanations apply a conformal-anomaly-detection-inspired filter to candidate perturbations, retaining a representative candidate only when its empirical guard score meets a user-chosen threshold $\varepsilon$. To facilitate this, we replace the standard binary discretiser with a multi-bin discretiser, generating interval-valued conditions alongside one-sided thresholds. Emitted perturbations are then scored using the standard Calibrated Explanations uncertainty-aware backend. In a synthetic constraint setting, we demonstrate that guarded factual explanations reduce the fraction of emitted rules whose representative values violate a known domain constraint from approximately 2.5% to near 0.0% at $\varepsilon$ = 0.2. Crucially, we isolate this effect from discretisation granularity alone by comparing against an unpruned multi-bin (no-guard) baseline. We frame guarded explanations as a representative-level plausibility filter, where the guard demonstrably reduces calibration-unsupported representative candidates in emitted rules. The guard score is a conformal-style empirical rank, not a standard conformal p-value, and $\varepsilon$ has no finite-sample error-rate guarantee for constructed perturbations.} }
Endnote
%0 Conference Paper %T Guarded Explanations: Conformal-Style Filtering for Distribution-Aware Rule Conditions %A Tuwe Löfström %A Anders Hjort %A Helena Löfström %B Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications %C Proceedings of Machine Learning Research %D 2026 %E Ernst Ahlberg %E Ulf Johansson %E Henrik Boström %E Alberto Carlevaro %E Johan Hallberg Szabadváry %E Lars Carlsson %F pmlr-v329-lofstrom26a %I PMLR %P 422--438 %U https://proceedings.mlr.press/v329/lofstrom26a.html %V 329 %X In high-stakes decision support, post-hoc explanation methods are often used to estimate feature influence by perturbing inputs. However, a fundamental limitation of this approach is that candidate perturbations frequently fall outside the support of the calibration data, yielding unsupported and potentially misleading explanations. We introduce guarded explanations, an extension of the Calibrated Explanations framework. Guarded explanations apply a conformal-anomaly-detection-inspired filter to candidate perturbations, retaining a representative candidate only when its empirical guard score meets a user-chosen threshold $\varepsilon$. To facilitate this, we replace the standard binary discretiser with a multi-bin discretiser, generating interval-valued conditions alongside one-sided thresholds. Emitted perturbations are then scored using the standard Calibrated Explanations uncertainty-aware backend. In a synthetic constraint setting, we demonstrate that guarded factual explanations reduce the fraction of emitted rules whose representative values violate a known domain constraint from approximately 2.5% to near 0.0% at $\varepsilon$ = 0.2. Crucially, we isolate this effect from discretisation granularity alone by comparing against an unpruned multi-bin (no-guard) baseline. We frame guarded explanations as a representative-level plausibility filter, where the guard demonstrably reduces calibration-unsupported representative candidates in emitted rules. The guard score is a conformal-style empirical rank, not a standard conformal p-value, and $\varepsilon$ has no finite-sample error-rate guarantee for constructed perturbations.
APA
Löfström, T., Hjort, A. & Löfström, H.. (2026). Guarded Explanations: Conformal-Style Filtering for Distribution-Aware Rule Conditions. Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, in Proceedings of Machine Learning Research 329:422-438 Available from https://proceedings.mlr.press/v329/lofstrom26a.html.

Related Material