[edit]
Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:531-540, 2017.
Abstract
$\tilde{P}$ $\tilde{N}$ learning is the problem of binary classi- fication when training examples may be mis- labeled (flipped) uniformly with noise rate $\rho$1 for positive examples and $\rho$0 for negative ex- amples. We propose Rank Pruning (RP) to solve $\tilde{P}$ $\tilde{N}$ learning and the open problem of es- timating the noise rates. Unlike prior solutions, RP is efficient and general, requiring O(T) for any unrestricted choice of probabilistic classi- fier with T fitting time. We prove RP achieves consistent noise estimation and equivalent ex- pected risk as learning with uncorrupted labels in ideal conditions, and derive closed-form so- lutions when conditions are non-ideal. RP achieves state-of-the-art noise estimation and F1, error, and AUC-PR for both MNIST and CIFAR datasets, regardless of the amount of noise. To highlight, RP with a CNN classifier can predict if an MNIST digit is a one or not with only 0.25% error, and 0.46% error across all digits, even when 50% of positive examples are mislabeled and 50% of observed positive labels are mislabeled negative examples.