Learning to Predict from Crowdsourced Data

Wei Bi, Liwei Wang UIUC, James Kwok, Zhuowen Tu UCSD
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:342-351, 2014.

Abstract

Crowdsourcing services like Amazon’s Mechan- ical Turk have facilitated and greatly expedited the manual labeling process from a large number of human workers. However, spammers are often unavoidable and the crowdsourced labels can be very noisy. In this paper, we explicitly account for four sources for a noisy crowdsourced label: worker’s dedication to the task, his/her expertise, his/her default labeling judgement, and sample difficulty. A novel mixture model is employed for worker annotations, which learns a prediction model directly from samples to labels for effi- cient out-of-sample testing. Experiments on both simulated and real-world crowdsourced data sets show that the proposed method achieves signifi- cant improvements over the state-of-the-art.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-bi14a, title = {Learning to Predict from Crowdsourced Data}, author = {Bi, Wei and UIUC, Liwei Wang and Kwok, James and UCSD, Zhuowen Tu}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {342--351}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/bi14a/bi14a.pdf}, url = {https://proceedings.mlr.press/r12/bi14a.html}, abstract = {Crowdsourcing services like Amazon’s Mechan- ical Turk have facilitated and greatly expedited the manual labeling process from a large number of human workers. However, spammers are often unavoidable and the crowdsourced labels can be very noisy. In this paper, we explicitly account for four sources for a noisy crowdsourced label: worker’s dedication to the task, his/her expertise, his/her default labeling judgement, and sample difficulty. A novel mixture model is employed for worker annotations, which learns a prediction model directly from samples to labels for effi- cient out-of-sample testing. Experiments on both simulated and real-world crowdsourced data sets show that the proposed method achieves signifi- cant improvements over the state-of-the-art.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Learning to Predict from Crowdsourced Data %A Wei Bi %A Liwei Wang UIUC %A James Kwok %A Zhuowen Tu UCSD %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-bi14a %I PMLR %P 342--351 %U https://proceedings.mlr.press/r12/bi14a.html %V R12 %X Crowdsourcing services like Amazon’s Mechan- ical Turk have facilitated and greatly expedited the manual labeling process from a large number of human workers. However, spammers are often unavoidable and the crowdsourced labels can be very noisy. In this paper, we explicitly account for four sources for a noisy crowdsourced label: worker’s dedication to the task, his/her expertise, his/her default labeling judgement, and sample difficulty. A novel mixture model is employed for worker annotations, which learns a prediction model directly from samples to labels for effi- cient out-of-sample testing. Experiments on both simulated and real-world crowdsourced data sets show that the proposed method achieves signifi- cant improvements over the state-of-the-art. %Z Reissued by PMLR on 04 October 2026.
APA
Bi, W., UIUC, L.W., Kwok, J. & UCSD, Z.T.. (2014). Learning to Predict from Crowdsourced Data. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:342-351 Available from https://proceedings.mlr.press/r12/bi14a.html. Reissued by PMLR on 04 October 2026.

Related Material