Is Supervised Learning Really That Different From Unsupervised?

Oskar Allerbo, Thomas B. Schön
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4663-4671, 2026.

Abstract

We demonstrate how supervised learning can be decomposed into a two-stage procedure, where (1) all model parameters are selected in an unsupervised manner, and (2) the outputs y are added to the model, without changing the parameter values. This is achieved by a new model selection criterion that - in contrast to cross-validation - can be used also without access to y. For linear ridge regression, we bound the asymptotic out-of-sample risk of our method in terms of the optimal asymptotic risk. We also demonstrate that versions of linear and kernel ridge regression, smoothing splines, k-nearest neighbors, random forests, and neural networks, trained without access to y, perform similarly to their standard y-based counterparts. Hence, our results suggest that the difference between supervised and unsupervised learning is less fundamental than it may appear.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-allerbo26a, title = { Is Supervised Learning Really That Different From Unsupervised? }, author = {Allerbo, Oskar and Sch{\"o}n, Thomas B.}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4663--4671}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/allerbo26a/allerbo26a.pdf}, url = {https://proceedings.mlr.press/v300/allerbo26a.html}, abstract = { We demonstrate how supervised learning can be decomposed into a two-stage procedure, where (1) all model parameters are selected in an unsupervised manner, and (2) the outputs y are added to the model, without changing the parameter values. This is achieved by a new model selection criterion that - in contrast to cross-validation - can be used also without access to y. For linear ridge regression, we bound the asymptotic out-of-sample risk of our method in terms of the optimal asymptotic risk. We also demonstrate that versions of linear and kernel ridge regression, smoothing splines, k-nearest neighbors, random forests, and neural networks, trained without access to y, perform similarly to their standard y-based counterparts. Hence, our results suggest that the difference between supervised and unsupervised learning is less fundamental than it may appear. } }
Endnote
%0 Conference Paper %T Is Supervised Learning Really That Different From Unsupervised? %A Oskar Allerbo %A Thomas B. Schön %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-allerbo26a %I PMLR %P 4663--4671 %U https://proceedings.mlr.press/v300/allerbo26a.html %V 300 %X We demonstrate how supervised learning can be decomposed into a two-stage procedure, where (1) all model parameters are selected in an unsupervised manner, and (2) the outputs y are added to the model, without changing the parameter values. This is achieved by a new model selection criterion that - in contrast to cross-validation - can be used also without access to y. For linear ridge regression, we bound the asymptotic out-of-sample risk of our method in terms of the optimal asymptotic risk. We also demonstrate that versions of linear and kernel ridge regression, smoothing splines, k-nearest neighbors, random forests, and neural networks, trained without access to y, perform similarly to their standard y-based counterparts. Hence, our results suggest that the difference between supervised and unsupervised learning is less fundamental than it may appear.
APA
Allerbo, O. & Schön, T.B.. (2026). Is Supervised Learning Really That Different From Unsupervised? . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4663-4671 Available from https://proceedings.mlr.press/v300/allerbo26a.html.

Related Material