A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools

Gerardo Flores, Alyssa Hasegawa Smith, Abigail E. Schiff, Julia Fukuyama, Ashia C. Wilson
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:505-513, 2026.

Abstract

Machine learning-supported decisions, such as ordering diagnostic tests or determining preventive custody, often rely on binary classification from probabilistic forecasts. A consequentialist perspective, long emphasized in decision theory, favors evaluation methods that reflect the quality of such forecasts under threshold uncertainty and varying prevalence, notably Brier scores and log loss. However, our empirical review of practices at major ML venues (ICML, FAccT, CHIL) reveals a dominant reliance on accuracy and AUC-ROC. To address this disconnect, we introduce a decision-theoretic framework mapping evaluation metrics to their appropriate use cases, along with a practical Python package, \texttt{briertools}, designed to make proper scoring rules more usable in real-world settings. Specifically, we implement a bounded-threshold variant of the Brier score and log loss that restricts evaluation to a practitioner-specified range of plausible cost ratios, rather than averaging over the full unit interval. We further contribute a theoretical reconciliation between the Brier score and decision curve analysis, directly addressing a longstanding critique by Assel et al (2017) regarding the clinical utility of proper scoring rules.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-flores26a, title = { A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools }, author = {Flores, Gerardo and Smith, Alyssa Hasegawa and Schiff, Abigail E. and Fukuyama, Julia and Wilson, Ashia C.}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {505--513}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/flores26a/flores26a.pdf}, url = {https://proceedings.mlr.press/v300/flores26a.html}, abstract = { Machine learning-supported decisions, such as ordering diagnostic tests or determining preventive custody, often rely on binary classification from probabilistic forecasts. A consequentialist perspective, long emphasized in decision theory, favors evaluation methods that reflect the quality of such forecasts under threshold uncertainty and varying prevalence, notably Brier scores and log loss. However, our empirical review of practices at major ML venues (ICML, FAccT, CHIL) reveals a dominant reliance on accuracy and AUC-ROC. To address this disconnect, we introduce a decision-theoretic framework mapping evaluation metrics to their appropriate use cases, along with a practical Python package, \texttt{briertools}, designed to make proper scoring rules more usable in real-world settings. Specifically, we implement a bounded-threshold variant of the Brier score and log loss that restricts evaluation to a practitioner-specified range of plausible cost ratios, rather than averaging over the full unit interval. We further contribute a theoretical reconciliation between the Brier score and decision curve analysis, directly addressing a longstanding critique by Assel et al (2017) regarding the clinical utility of proper scoring rules. } }
Endnote
%0 Conference Paper %T A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools %A Gerardo Flores %A Alyssa Hasegawa Smith %A Abigail E. Schiff %A Julia Fukuyama %A Ashia C. Wilson %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-flores26a %I PMLR %P 505--513 %U https://proceedings.mlr.press/v300/flores26a.html %V 300 %X Machine learning-supported decisions, such as ordering diagnostic tests or determining preventive custody, often rely on binary classification from probabilistic forecasts. A consequentialist perspective, long emphasized in decision theory, favors evaluation methods that reflect the quality of such forecasts under threshold uncertainty and varying prevalence, notably Brier scores and log loss. However, our empirical review of practices at major ML venues (ICML, FAccT, CHIL) reveals a dominant reliance on accuracy and AUC-ROC. To address this disconnect, we introduce a decision-theoretic framework mapping evaluation metrics to their appropriate use cases, along with a practical Python package, \texttt{briertools}, designed to make proper scoring rules more usable in real-world settings. Specifically, we implement a bounded-threshold variant of the Brier score and log loss that restricts evaluation to a practitioner-specified range of plausible cost ratios, rather than averaging over the full unit interval. We further contribute a theoretical reconciliation between the Brier score and decision curve analysis, directly addressing a longstanding critique by Assel et al (2017) regarding the clinical utility of proper scoring rules.
APA
Flores, G., Smith, A.H., Schiff, A.E., Fukuyama, J. & Wilson, A.C.. (2026). A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:505-513 Available from https://proceedings.mlr.press/v300/flores26a.html.

Related Material