[edit]
Performance Estimation in Hybrid Interpretable Models with Venn Predictors
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:439-462, 2026.
Abstract
The design of hybrid interpretable models has recently emerged as a promising paradigm for the development of explainable artificial intelligence. This modeling framework relies on a collaborative scheme in which black-box and interpretable models are paired and cooperate to produce a final prediction. Intuitively, it assumes that there are some regions of the feature space where a black-box classifier can be replaced by an interpretable model without losing predictive performance. For hybrid interpretable models to be useful in practice, users should have statistical guarantees on their performance for a given transparency level (i.e., the fraction of samples delegated to the interpretable classifier). In this work, we propose the use of Venn prediction to construct reliable accuracy estimators for hybrid interpretable models in the absence of ground truth labels. In particular, we derive estimators for the marginal accuracy (i.e., for the hybrid model as a single predictor) and the component-conditional accuracy (i.e., for each of the internal components of the hybrid model). We demonstrate the benefits of our estimation methodology, consistently outperforming alternative approaches for six different datasets, with the conditional Venn predictor providing the most reliable component-conditional estimates.