[edit]
Identifying Labeling Mechanism in Positive–Unlabeled Learning under Unknown Class Prior
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2119-2136, 2026.
Abstract
Positive–Unlabeled ({PU}) learning critically depends on assumptions about how positive instances are labeled, yet the labeling mechanism is typically unknown in practice. Misspecification of the labeling mechanism can therefore lead to systematic bias and unstable learning. We formulate labeling mechanism identification as a statistical inference problem under an unknown class prior using only positive and unlabeled data. We propose a two-stage testing framework that first constructs a family of empirical positive sets through false discovery rate ({FDR})-controlled multiple testing at multiple {FDR} levels, and then performs bootstrap-based selected completely at random ({SCAR}) consistency tests on each empirical positive set, whose evidence is aggregated via a Bonferroni correction to produce a global decision.We establish that each empirical positive set admits provably controlled contamination, and further show that the Bonferroni-aggregated bootstrap test provides valid inference. Experiments on synthetic and real-world datasets demonstrate that the proposed framework reliably distinguishes {SCAR} from selected at random (SAR) mechanisms and improves the robustness of downstream {PU} learning.