Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels

Curtis G. Northcutt, Tailin Wu, Isaac L. Chuang
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:531-540, 2017.

Abstract

$\tilde{P}$ $\tilde{N}$ learning is the problem of binary classi- fication when training examples may be mis- labeled (flipped) uniformly with noise rate $\rho$1 for positive examples and $\rho$0 for negative ex- amples. We propose Rank Pruning (RP) to solve $\tilde{P}$ $\tilde{N}$ learning and the open problem of es- timating the noise rates. Unlike prior solutions, RP is efficient and general, requiring O(T) for any unrestricted choice of probabilistic classi- fier with T fitting time. We prove RP achieves consistent noise estimation and equivalent ex- pected risk as learning with uncorrupted labels in ideal conditions, and derive closed-form so- lutions when conditions are non-ideal. RP achieves state-of-the-art noise estimation and F1, error, and AUC-PR for both MNIST and CIFAR datasets, regardless of the amount of noise. To highlight, RP with a CNN classifier can predict if an MNIST digit is a one or not with only 0.25% error, and 0.46% error across all digits, even when 50% of positive examples are mislabeled and 50% of observed positive labels are mislabeled negative examples.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR15-northcutt17a, title = {Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels}, author = {Northcutt, Curtis G. and Wu, Tailin and Chuang, Isaac L.}, booktitle = {Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence}, pages = {531--540}, year = {2017}, editor = {Elidan, Gal and Kersting, Kristian}, volume = {R15}, series = {Proceedings of Machine Learning Research}, month = {11--15 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r15/main/assets/northcutt17a/northcutt17a.pdf}, url = {https://proceedings.mlr.press/r15/northcutt17a.html}, abstract = {$\tilde{P}$ $\tilde{N}$ learning is the problem of binary classi- fication when training examples may be mis- labeled (flipped) uniformly with noise rate $\rho$1 for positive examples and $\rho$0 for negative ex- amples. We propose Rank Pruning (RP) to solve $\tilde{P}$ $\tilde{N}$ learning and the open problem of es- timating the noise rates. Unlike prior solutions, RP is efficient and general, requiring O(T) for any unrestricted choice of probabilistic classi- fier with T fitting time. We prove RP achieves consistent noise estimation and equivalent ex- pected risk as learning with uncorrupted labels in ideal conditions, and derive closed-form so- lutions when conditions are non-ideal. RP achieves state-of-the-art noise estimation and F1, error, and AUC-PR for both MNIST and CIFAR datasets, regardless of the amount of noise. To highlight, RP with a CNN classifier can predict if an MNIST digit is a one or not with only 0.25% error, and 0.46% error across all digits, even when 50% of positive examples are mislabeled and 50% of observed positive labels are mislabeled negative examples.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels %A Curtis G. Northcutt %A Tailin Wu %A Isaac L. Chuang %B Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2017 %E Gal Elidan %E Kristian Kersting %F pmlr-vR15-northcutt17a %I PMLR %P 531--540 %U https://proceedings.mlr.press/r15/northcutt17a.html %V R15 %X $\tilde{P}$ $\tilde{N}$ learning is the problem of binary classi- fication when training examples may be mis- labeled (flipped) uniformly with noise rate $\rho$1 for positive examples and $\rho$0 for negative ex- amples. We propose Rank Pruning (RP) to solve $\tilde{P}$ $\tilde{N}$ learning and the open problem of es- timating the noise rates. Unlike prior solutions, RP is efficient and general, requiring O(T) for any unrestricted choice of probabilistic classi- fier with T fitting time. We prove RP achieves consistent noise estimation and equivalent ex- pected risk as learning with uncorrupted labels in ideal conditions, and derive closed-form so- lutions when conditions are non-ideal. RP achieves state-of-the-art noise estimation and F1, error, and AUC-PR for both MNIST and CIFAR datasets, regardless of the amount of noise. To highlight, RP with a CNN classifier can predict if an MNIST digit is a one or not with only 0.25% error, and 0.46% error across all digits, even when 50% of positive examples are mislabeled and 50% of observed positive labels are mislabeled negative examples. %Z Reissued by PMLR on 04 October 2026.
APA
Northcutt, C.G., Wu, T. & Chuang, I.L.. (2017). Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels. Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R15:531-540 Available from https://proceedings.mlr.press/r15/northcutt17a.html. Reissued by PMLR on 04 October 2026.

Related Material