Too Good to Conform: The Inverse Quality Paradox in Counterfactual Filtering

Fatima Rabia Yapicioglu, Abhishek Srinivasan, Juan Carlos Andresen, Henrik Boström
Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR 329:1073-1075, 2026.

Abstract

As machine learning models enter high-stakes decision-making, explanations must be interpretable, reliable, and uncertainty-aware (Löfström et al., 2024, 2026). Counterfactuals (CFs) are a prominent explanation form, but their plausibility remains difficult to guarantee (Guidotti, 2024). Conformal prediction (Vovk et al., 2022) addresses this through generate-then-filter, which validates candidates post-hoc (Bates et al., 2023; Maalej et al., 2025; Adams et al., 2025), or filter-then-generate, which incorporates conformal constraints into generation (Altmeyer et al., 2024; Bilkhoo et al., 2025). We study the former. Although nonconformity-measure (NCM) design is actively studied in other conformal settings (Narteni et al., 2026), its role as a structural failure point in counterfactual filtering remains unexplored. We show that an NCM assigning greater nonconformity to candidates closer to xi can systematically reject the most proximal valid CFs — as generator quality (proximity to xi among candidates achieving the desired prediction change) improves, nonconformity increases, producing an inverse quality paradox. Our contributions are to: (i) formalise this failure for the reciprocal-distance score Aprox; (ii) show that it can reject reference-supported candidates while accepting candidates outside the data-supported region; and (iii) demonstrate that a score AR anchored to a fixed in-distribution reference set removes this dependence on xi and retains standard conformal validity under exchangeability.

Cite this Paper


BibTeX
@InProceedings{pmlr-v329-rabia-yapicioglu26a, title = {Too Good to Conform: The Inverse Quality Paradox in Counterfactual Filtering}, author = {Rabia Yapicioglu, Fatima and Srinivasan, Abhishek and Carlos Andresen, Juan and Bostr{\"o}m, Henrik}, booktitle = {Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications}, pages = {1073--1075}, year = {2026}, editor = {Ahlberg, Ernst and Johansson, Ulf and Boström, Henrik and Carlevaro, Alberto and Hallberg Szabadváry, Johan and Carlsson, Lars}, volume = {329}, series = {Proceedings of Machine Learning Research}, month = {02--04 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v329/main/assets/rabia-yapicioglu26a/rabia-yapicioglu26a.pdf}, url = {https://proceedings.mlr.press/v329/rabia-yapicioglu26a.html}, abstract = {As machine learning models enter high-stakes decision-making, explanations must be interpretable, reliable, and uncertainty-aware (Löfström et al., 2024, 2026). Counterfactuals (CFs) are a prominent explanation form, but their plausibility remains difficult to guarantee (Guidotti, 2024). Conformal prediction (Vovk et al., 2022) addresses this through generate-then-filter, which validates candidates post-hoc (Bates et al., 2023; Maalej et al., 2025; Adams et al., 2025), or filter-then-generate, which incorporates conformal constraints into generation (Altmeyer et al., 2024; Bilkhoo et al., 2025). We study the former. Although nonconformity-measure (NCM) design is actively studied in other conformal settings (Narteni et al., 2026), its role as a structural failure point in counterfactual filtering remains unexplored. We show that an NCM assigning greater nonconformity to candidates closer to xi can systematically reject the most proximal valid CFs — as generator quality (proximity to xi among candidates achieving the desired prediction change) improves, nonconformity increases, producing an inverse quality paradox. Our contributions are to: (i) formalise this failure for the reciprocal-distance score Aprox; (ii) show that it can reject reference-supported candidates while accepting candidates outside the data-supported region; and (iii) demonstrate that a score AR anchored to a fixed in-distribution reference set removes this dependence on xi and retains standard conformal validity under exchangeability.} }
Endnote
%0 Conference Paper %T Too Good to Conform: The Inverse Quality Paradox in Counterfactual Filtering %A Fatima Rabia Yapicioglu %A Abhishek Srinivasan %A Juan Carlos Andresen %A Henrik Boström %B Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications %C Proceedings of Machine Learning Research %D 2026 %E Ernst Ahlberg %E Ulf Johansson %E Henrik Boström %E Alberto Carlevaro %E Johan Hallberg Szabadváry %E Lars Carlsson %F pmlr-v329-rabia-yapicioglu26a %I PMLR %P 1073--1075 %U https://proceedings.mlr.press/v329/rabia-yapicioglu26a.html %V 329 %X As machine learning models enter high-stakes decision-making, explanations must be interpretable, reliable, and uncertainty-aware (Löfström et al., 2024, 2026). Counterfactuals (CFs) are a prominent explanation form, but their plausibility remains difficult to guarantee (Guidotti, 2024). Conformal prediction (Vovk et al., 2022) addresses this through generate-then-filter, which validates candidates post-hoc (Bates et al., 2023; Maalej et al., 2025; Adams et al., 2025), or filter-then-generate, which incorporates conformal constraints into generation (Altmeyer et al., 2024; Bilkhoo et al., 2025). We study the former. Although nonconformity-measure (NCM) design is actively studied in other conformal settings (Narteni et al., 2026), its role as a structural failure point in counterfactual filtering remains unexplored. We show that an NCM assigning greater nonconformity to candidates closer to xi can systematically reject the most proximal valid CFs — as generator quality (proximity to xi among candidates achieving the desired prediction change) improves, nonconformity increases, producing an inverse quality paradox. Our contributions are to: (i) formalise this failure for the reciprocal-distance score Aprox; (ii) show that it can reject reference-supported candidates while accepting candidates outside the data-supported region; and (iii) demonstrate that a score AR anchored to a fixed in-distribution reference set removes this dependence on xi and retains standard conformal validity under exchangeability.
APA
Rabia Yapicioglu, F., Srinivasan, A., Carlos Andresen, J. & Boström, H.. (2026). Too Good to Conform: The Inverse Quality Paradox in Counterfactual Filtering. Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications, in Proceedings of Machine Learning Research 329:1073-1075 Available from https://proceedings.mlr.press/v329/rabia-yapicioglu26a.html.

Related Material