Positive-Unlabeled Regression: learning from partially labeled quantitative outcomes

Paweł Teisseyre, Jan Mielniczuk
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:6677-6700, 2026.

Abstract

We study a novel machine learning problem: Positive-Unlabeled Regression (PUR), which extends the classical Positive-Unlabeled ({PU}) learning framework to regression setting with a non-negative, discrete target variable. This setting arises naturally in applications where only some positive quantitative outcomes are reported, while the absence of a label may either indicate a true zero outcome or an unreported positive value. Applications include predicting the number of diseases a patient may have, the number of complications following an illness, or the advancement stage of a disease. We formalize the PUR problem and highlight the limitations of naive approaches that either use reported target variable or discard unlabeled data. To account for the inherent bias in such strategies, we propose two principled methods. The first is based on calibration of regression estimates using posterior probabilities from classical {PU} learning. The second builds on an empirical risk minimization framework, restating the target risk as a weighted function dependent on the instance-specific propensity score. We demonstrate both theoretically and empirically that the proposed approaches yield improved performance over standard baselines.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-teisseyre26a, title = {Positive-Unlabeled Regression: learning from partially labeled quantitative outcomes}, author = {Teisseyre, Pawe{\l} and Mielniczuk, Jan}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {6677--6700}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/teisseyre26a/teisseyre26a.pdf}, url = {https://proceedings.mlr.press/v337/teisseyre26a.html}, abstract = {We study a novel machine learning problem: Positive-Unlabeled Regression (PUR), which extends the classical Positive-Unlabeled ({PU}) learning framework to regression setting with a non-negative, discrete target variable. This setting arises naturally in applications where only some positive quantitative outcomes are reported, while the absence of a label may either indicate a true zero outcome or an unreported positive value. Applications include predicting the number of diseases a patient may have, the number of complications following an illness, or the advancement stage of a disease. We formalize the PUR problem and highlight the limitations of naive approaches that either use reported target variable or discard unlabeled data. To account for the inherent bias in such strategies, we propose two principled methods. The first is based on calibration of regression estimates using posterior probabilities from classical {PU} learning. The second builds on an empirical risk minimization framework, restating the target risk as a weighted function dependent on the instance-specific propensity score. We demonstrate both theoretically and empirically that the proposed approaches yield improved performance over standard baselines.} }
Endnote
%0 Conference Paper %T Positive-Unlabeled Regression: learning from partially labeled quantitative outcomes %A Paweł Teisseyre %A Jan Mielniczuk %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-teisseyre26a %I PMLR %P 6677--6700 %U https://proceedings.mlr.press/v337/teisseyre26a.html %V 337 %X We study a novel machine learning problem: Positive-Unlabeled Regression (PUR), which extends the classical Positive-Unlabeled ({PU}) learning framework to regression setting with a non-negative, discrete target variable. This setting arises naturally in applications where only some positive quantitative outcomes are reported, while the absence of a label may either indicate a true zero outcome or an unreported positive value. Applications include predicting the number of diseases a patient may have, the number of complications following an illness, or the advancement stage of a disease. We formalize the PUR problem and highlight the limitations of naive approaches that either use reported target variable or discard unlabeled data. To account for the inherent bias in such strategies, we propose two principled methods. The first is based on calibration of regression estimates using posterior probabilities from classical {PU} learning. The second builds on an empirical risk minimization framework, restating the target risk as a weighted function dependent on the instance-specific propensity score. We demonstrate both theoretically and empirically that the proposed approaches yield improved performance over standard baselines.
APA
Teisseyre, P. & Mielniczuk, J.. (2026). Positive-Unlabeled Regression: learning from partially labeled quantitative outcomes. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:6677-6700 Available from https://proceedings.mlr.press/v337/teisseyre26a.html.

Related Material