Identifying Labeling Mechanism in Positive–Unlabeled Learning under Unknown Class Prior

Siying He, Weijuan Liang, Jiatong Liu
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2119-2136, 2026.

Abstract

Positive–Unlabeled ({PU}) learning critically depends on assumptions about how positive instances are labeled, yet the labeling mechanism is typically unknown in practice. Misspecification of the labeling mechanism can therefore lead to systematic bias and unstable learning. We formulate labeling mechanism identification as a statistical inference problem under an unknown class prior using only positive and unlabeled data. We propose a two-stage testing framework that first constructs a family of empirical positive sets through false discovery rate ({FDR})-controlled multiple testing at multiple {FDR} levels, and then performs bootstrap-based selected completely at random ({SCAR}) consistency tests on each empirical positive set, whose evidence is aggregated via a Bonferroni correction to produce a global decision.We establish that each empirical positive set admits provably controlled contamination, and further show that the Bonferroni-aggregated bootstrap test provides valid inference. Experiments on synthetic and real-world datasets demonstrate that the proposed framework reliably distinguishes {SCAR} from selected at random (SAR) mechanisms and improves the robustness of downstream {PU} learning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-he26a, title = {Identifying Labeling Mechanism in Positive–Unlabeled Learning under Unknown Class Prior}, author = {He, Siying and Liang, Weijuan and Liu, Jiatong}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2119--2136}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/he26a/he26a.pdf}, url = {https://proceedings.mlr.press/v337/he26a.html}, abstract = {Positive–Unlabeled ({PU}) learning critically depends on assumptions about how positive instances are labeled, yet the labeling mechanism is typically unknown in practice. Misspecification of the labeling mechanism can therefore lead to systematic bias and unstable learning. We formulate labeling mechanism identification as a statistical inference problem under an unknown class prior using only positive and unlabeled data. We propose a two-stage testing framework that first constructs a family of empirical positive sets through false discovery rate ({FDR})-controlled multiple testing at multiple {FDR} levels, and then performs bootstrap-based selected completely at random ({SCAR}) consistency tests on each empirical positive set, whose evidence is aggregated via a Bonferroni correction to produce a global decision.We establish that each empirical positive set admits provably controlled contamination, and further show that the Bonferroni-aggregated bootstrap test provides valid inference. Experiments on synthetic and real-world datasets demonstrate that the proposed framework reliably distinguishes {SCAR} from selected at random (SAR) mechanisms and improves the robustness of downstream {PU} learning.} }
Endnote
%0 Conference Paper %T Identifying Labeling Mechanism in Positive–Unlabeled Learning under Unknown Class Prior %A Siying He %A Weijuan Liang %A Jiatong Liu %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-he26a %I PMLR %P 2119--2136 %U https://proceedings.mlr.press/v337/he26a.html %V 337 %X Positive–Unlabeled ({PU}) learning critically depends on assumptions about how positive instances are labeled, yet the labeling mechanism is typically unknown in practice. Misspecification of the labeling mechanism can therefore lead to systematic bias and unstable learning. We formulate labeling mechanism identification as a statistical inference problem under an unknown class prior using only positive and unlabeled data. We propose a two-stage testing framework that first constructs a family of empirical positive sets through false discovery rate ({FDR})-controlled multiple testing at multiple {FDR} levels, and then performs bootstrap-based selected completely at random ({SCAR}) consistency tests on each empirical positive set, whose evidence is aggregated via a Bonferroni correction to produce a global decision.We establish that each empirical positive set admits provably controlled contamination, and further show that the Bonferroni-aggregated bootstrap test provides valid inference. Experiments on synthetic and real-world datasets demonstrate that the proposed framework reliably distinguishes {SCAR} from selected at random (SAR) mechanisms and improves the robustness of downstream {PU} learning.
APA
He, S., Liang, W. & Liu, J.. (2026). Identifying Labeling Mechanism in Positive–Unlabeled Learning under Unknown Class Prior. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2119-2136 Available from https://proceedings.mlr.press/v337/he26a.html.

Related Material