Semi-supervised Learning by Modeling Multiple-Annotator Expertise

Yan Yan, Romer Rosales, Glenn Fung, Jennifer Dy
Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, PMLR R8:673-681, 2010.

Abstract

Learning algorithms normally assume that there is at most one annotation or label per data point. However, in some scenarios, such as medical di- agnosis and on-line collaboration, multiple anno- tations may be available. In either case, obtain- ing labels for data points can be expensive and time-consuming (in some circumstances ground- truth may not exist). Semi-supervised learning approaches have shown that utilizing the unla- beled data is often beneficial in these cases. This paper presents a probabilistic semi-supervised model and algorithm that allows for learning from both unlabeled and labeled data in the pres- ence of multiple annotators. We assume that it is known what annotator labeled which data points. The proposed approach produces anno- tator models that allow us to provide (1) esti- mates of the true label and (2) annotator variable expertise for both labeled and unlabeled data. We provide numerical comparisons under vari- ous scenarios and with respect to standard semi- supervised learning. Experiments showed that the presented approach provides clear advantages over multi-annotator methods that do not use the unlabeled data and over methods that do not use multi-labeler information.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR8-yan10a, title = {Semi-supervised Learning by Modeling Multiple-Annotator Expertise}, author = {Yan, Yan and Rosales, Romer and Fung, Glenn and Dy, Jennifer}, booktitle = {Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence}, pages = {673--681}, year = {2010}, editor = {Grünwald, Peter and Spirtes, Peter}, volume = {R8}, series = {Proceedings of Machine Learning Research}, month = {08--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r8/main/assets/yan10a/yan10a.pdf}, url = {https://proceedings.mlr.press/r8/yan10a.html}, abstract = {Learning algorithms normally assume that there is at most one annotation or label per data point. However, in some scenarios, such as medical di- agnosis and on-line collaboration, multiple anno- tations may be available. In either case, obtain- ing labels for data points can be expensive and time-consuming (in some circumstances ground- truth may not exist). Semi-supervised learning approaches have shown that utilizing the unla- beled data is often beneficial in these cases. This paper presents a probabilistic semi-supervised model and algorithm that allows for learning from both unlabeled and labeled data in the pres- ence of multiple annotators. We assume that it is known what annotator labeled which data points. The proposed approach produces anno- tator models that allow us to provide (1) esti- mates of the true label and (2) annotator variable expertise for both labeled and unlabeled data. We provide numerical comparisons under vari- ous scenarios and with respect to standard semi- supervised learning. Experiments showed that the presented approach provides clear advantages over multi-annotator methods that do not use the unlabeled data and over methods that do not use multi-labeler information.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Semi-supervised Learning by Modeling Multiple-Annotator Expertise %A Yan Yan %A Romer Rosales %A Glenn Fung %A Jennifer Dy %B Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2010 %E Peter Grünwald %E Peter Spirtes %F pmlr-vR8-yan10a %I PMLR %P 673--681 %U https://proceedings.mlr.press/r8/yan10a.html %V R8 %X Learning algorithms normally assume that there is at most one annotation or label per data point. However, in some scenarios, such as medical di- agnosis and on-line collaboration, multiple anno- tations may be available. In either case, obtain- ing labels for data points can be expensive and time-consuming (in some circumstances ground- truth may not exist). Semi-supervised learning approaches have shown that utilizing the unla- beled data is often beneficial in these cases. This paper presents a probabilistic semi-supervised model and algorithm that allows for learning from both unlabeled and labeled data in the pres- ence of multiple annotators. We assume that it is known what annotator labeled which data points. The proposed approach produces anno- tator models that allow us to provide (1) esti- mates of the true label and (2) annotator variable expertise for both labeled and unlabeled data. We provide numerical comparisons under vari- ous scenarios and with respect to standard semi- supervised learning. Experiments showed that the presented approach provides clear advantages over multi-annotator methods that do not use the unlabeled data and over methods that do not use multi-labeler information. %Z Reissued by PMLR on 04 October 2026.
APA
Yan, Y., Rosales, R., Fung, G. & Dy, J.. (2010). Semi-supervised Learning by Modeling Multiple-Annotator Expertise. Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R8:673-681 Available from https://proceedings.mlr.press/r8/yan10a.html. Reissued by PMLR on 04 October 2026.

Related Material