[edit]
Explaining Conformal Prediction: Diagnosing Reliability through Feature-Conditioned Outcomes and p-Value Margins
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:509-528, 2026.
Abstract
Conformal prediction produces prediction sets with finite-sample coverage guarantees, and its practical utility for decision-making grows when the reliability of its predictions can be explained at the feature level. Yet, the intersection of conformal prediction and explainability remains relatively underexplored, leaving practitioners without tools to interpret how input features shape prediction reliability. We introduce a framework for explaining the reliability of conformal classifiers at the feature level using two tools. First, Feature-Conditioned Outcome Distribution (FCOD) plots visualize how correct, incorrect, and ambiguous conformal prediction set outcomes vary across feature ranges, identifying regions of the feature space associated with different empirical behaviors in terms of correctness and uncertainty of the predictions. Second, we propose two metrics derived from class-conditional conformal p-values: the label-free evidence margin, which provides a signed contrast between competing classes without the true label, and the label-dependent correctness margin, which compares the conformal p-value of the true class with that of its strongest alternative. They are used as SHAP targets for local feature attribution and aggregated global analysis, enabling feature-level explanations of relative conformal evidence between competing classes and support for the true class.