[edit]
Semi-supervised Learning by Modeling Multiple-Annotator Expertise
Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, PMLR R8:673-681, 2010.
Abstract
Learning algorithms normally assume that there is at most one annotation or label per data point. However, in some scenarios, such as medical di- agnosis and on-line collaboration, multiple anno- tations may be available. In either case, obtain- ing labels for data points can be expensive and time-consuming (in some circumstances ground- truth may not exist). Semi-supervised learning approaches have shown that utilizing the unla- beled data is often beneficial in these cases. This paper presents a probabilistic semi-supervised model and algorithm that allows for learning from both unlabeled and labeled data in the pres- ence of multiple annotators. We assume that it is known what annotator labeled which data points. The proposed approach produces anno- tator models that allow us to provide (1) esti- mates of the true label and (2) annotator variable expertise for both labeled and unlabeled data. We provide numerical comparisons under vari- ous scenarios and with respect to standard semi- supervised learning. Experiments showed that the presented approach provides clear advantages over multi-annotator methods that do not use the unlabeled data and over methods that do not use multi-labeler information.