[edit]
Guarded Explanations: Conformal-Style Filtering for Distribution-Aware Rule Conditions
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:422-438, 2026.
Abstract
In high-stakes decision support, post-hoc explanation methods are often used to estimate feature influence by perturbing inputs. However, a fundamental limitation of this approach is that candidate perturbations frequently fall outside the support of the calibration data, yielding unsupported and potentially misleading explanations. We introduce guarded explanations, an extension of the Calibrated Explanations framework. Guarded explanations apply a conformal-anomaly-detection-inspired filter to candidate perturbations, retaining a representative candidate only when its empirical guard score meets a user-chosen threshold $\varepsilon$. To facilitate this, we replace the standard binary discretiser with a multi-bin discretiser, generating interval-valued conditions alongside one-sided thresholds. Emitted perturbations are then scored using the standard Calibrated Explanations uncertainty-aware backend. In a synthetic constraint setting, we demonstrate that guarded factual explanations reduce the fraction of emitted rules whose representative values violate a known domain constraint from approximately 2.5% to near 0.0% at $\varepsilon$ = 0.2. Crucially, we isolate this effect from discretisation granularity alone by comparing against an unpruned multi-bin (no-guard) baseline. We frame guarded explanations as a representative-level plausibility filter, where the guard demonstrably reduces calibration-unsupported representative candidates in emitted rules. The guard score is a conformal-style empirical rank, not a standard conformal p-value, and $\varepsilon$ has no finite-sample error-rate guarantee for constructed perturbations.