Estimating Accuracy from Unlabeled Data

Emmanouil Antonios Platanios Carnegie Mellon University, Avrim Blum Carnegie Mellon University, Tom Mitchell Carnegie Mellon University
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:226-235, 2014.

Abstract

We consider the question of how unlabeled data can be used to estimate the true accuracy of learned classifiers. This is an important question for any autonomous learning system that must es- timate its accuracy without supervision, and also when classifiers trained from one data distribu- tion must be applied to a new distribution (e.g., document classifiers trained on one text corpus are to be applied to a second corpus). We first show how to estimate error rates exactly from unlabeled data when given a collection of com- peting classifiers that make independent errors, based on the agreement rates between subsets of these classifiers. We further show that even when the competing classifiers do not make indepen- dent errors, both their accuracies and error de- pendencies can be estimated by making certain relaxed assumptions. Experiments on two data real-world data sets produce estimates within a few percent of the true accuracy, using solely un- labeled data. These results are of practical signif- icance in situations where labeled data is scarce and shed light on the more general question of how the consistency among multiple functions is related to their true accuracies.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-university14g, title = {Estimating Accuracy from Unlabeled Data}, author = {University, Emmanouil Antonios Platanios Carnegie Mellon and University, Avrim Blum Carnegie Mellon and University, Tom Mitchell Carnegie Mellon}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {226--235}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/university14g/university14g.pdf}, url = {https://proceedings.mlr.press/r12/university14g.html}, abstract = {We consider the question of how unlabeled data can be used to estimate the true accuracy of learned classifiers. This is an important question for any autonomous learning system that must es- timate its accuracy without supervision, and also when classifiers trained from one data distribu- tion must be applied to a new distribution (e.g., document classifiers trained on one text corpus are to be applied to a second corpus). We first show how to estimate error rates exactly from unlabeled data when given a collection of com- peting classifiers that make independent errors, based on the agreement rates between subsets of these classifiers. We further show that even when the competing classifiers do not make indepen- dent errors, both their accuracies and error de- pendencies can be estimated by making certain relaxed assumptions. Experiments on two data real-world data sets produce estimates within a few percent of the true accuracy, using solely un- labeled data. These results are of practical signif- icance in situations where labeled data is scarce and shed light on the more general question of how the consistency among multiple functions is related to their true accuracies.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Estimating Accuracy from Unlabeled Data %A Emmanouil Antonios Platanios Carnegie Mellon University %A Avrim Blum Carnegie Mellon University %A Tom Mitchell Carnegie Mellon University %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-university14g %I PMLR %P 226--235 %U https://proceedings.mlr.press/r12/university14g.html %V R12 %X We consider the question of how unlabeled data can be used to estimate the true accuracy of learned classifiers. This is an important question for any autonomous learning system that must es- timate its accuracy without supervision, and also when classifiers trained from one data distribu- tion must be applied to a new distribution (e.g., document classifiers trained on one text corpus are to be applied to a second corpus). We first show how to estimate error rates exactly from unlabeled data when given a collection of com- peting classifiers that make independent errors, based on the agreement rates between subsets of these classifiers. We further show that even when the competing classifiers do not make indepen- dent errors, both their accuracies and error de- pendencies can be estimated by making certain relaxed assumptions. Experiments on two data real-world data sets produce estimates within a few percent of the true accuracy, using solely un- labeled data. These results are of practical signif- icance in situations where labeled data is scarce and shed light on the more general question of how the consistency among multiple functions is related to their true accuracies. %Z Reissued by PMLR on 04 October 2026.
APA
University, E.A.P.C.M., University, A.B.C.M. & University, T.M.C.M.. (2014). Estimating Accuracy from Unlabeled Data. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:226-235 Available from https://proceedings.mlr.press/r12/university14g.html. Reissued by PMLR on 04 October 2026.

Related Material