On the Normalization of Confusion Matrices: Methods and Geometric Interpretations

Johan Erbani, Sonia Ben Mokhtar, Pierre-Edouard Portier, Elöd Egyed-Zsigmond, Diana Nurbakova
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4924-4932, 2026.

Abstract

The confusion matrix is a standard tool for evaluating classifiers, providing a detailed view of model errors. In heterogeneous settings, its entries are influenced by two main factors: class similarity, reflecting how easily the model confuses certain classes, and distribution bias, stemming from imbalanced training or test distributions. Because confusion matrix values jointly reflect both factors, it is difficult to disentangle their individual effects. To address this issue, we introduce bi-normalization via Iterative Proportional Fitting, a generalization of row and column normalization. Unlike standard approaches, this method recovers the underlying structure of class similarity. By disentangling error sources, it enables a more precise diagnosis of model behavior and facilitates classifier improvement. We further establish connections between normalization, importance sampling, and class representations in the model’s latent space, thus offering a clearer interpretation of normalization schemes. Our implementation is publicly available.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-erbani26a, title = { On the Normalization of Confusion Matrices: Methods and Geometric Interpretations }, author = {Erbani, Johan and Mokhtar, Sonia Ben and Portier, Pierre-Edouard and Egyed-Zsigmond, El{\"o}d and Nurbakova, Diana}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4924--4932}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/erbani26a/erbani26a.pdf}, url = {https://proceedings.mlr.press/v300/erbani26a.html}, abstract = { The confusion matrix is a standard tool for evaluating classifiers, providing a detailed view of model errors. In heterogeneous settings, its entries are influenced by two main factors: class similarity, reflecting how easily the model confuses certain classes, and distribution bias, stemming from imbalanced training or test distributions. Because confusion matrix values jointly reflect both factors, it is difficult to disentangle their individual effects. To address this issue, we introduce bi-normalization via Iterative Proportional Fitting, a generalization of row and column normalization. Unlike standard approaches, this method recovers the underlying structure of class similarity. By disentangling error sources, it enables a more precise diagnosis of model behavior and facilitates classifier improvement. We further establish connections between normalization, importance sampling, and class representations in the model’s latent space, thus offering a clearer interpretation of normalization schemes. Our implementation is publicly available. } }
Endnote
%0 Conference Paper %T On the Normalization of Confusion Matrices: Methods and Geometric Interpretations %A Johan Erbani %A Sonia Ben Mokhtar %A Pierre-Edouard Portier %A Elöd Egyed-Zsigmond %A Diana Nurbakova %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-erbani26a %I PMLR %P 4924--4932 %U https://proceedings.mlr.press/v300/erbani26a.html %V 300 %X The confusion matrix is a standard tool for evaluating classifiers, providing a detailed view of model errors. In heterogeneous settings, its entries are influenced by two main factors: class similarity, reflecting how easily the model confuses certain classes, and distribution bias, stemming from imbalanced training or test distributions. Because confusion matrix values jointly reflect both factors, it is difficult to disentangle their individual effects. To address this issue, we introduce bi-normalization via Iterative Proportional Fitting, a generalization of row and column normalization. Unlike standard approaches, this method recovers the underlying structure of class similarity. By disentangling error sources, it enables a more precise diagnosis of model behavior and facilitates classifier improvement. We further establish connections between normalization, importance sampling, and class representations in the model’s latent space, thus offering a clearer interpretation of normalization schemes. Our implementation is publicly available.
APA
Erbani, J., Mokhtar, S.B., Portier, P., Egyed-Zsigmond, E. & Nurbakova, D.. (2026). On the Normalization of Confusion Matrices: Methods and Geometric Interpretations . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4924-4932 Available from https://proceedings.mlr.press/v300/erbani26a.html.

Related Material