The Relative Instability of Model Comparison with Cross-validation

Alexandre Bayle, Lucas Janson, Lester Mackey
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:7144-7178, 2026.

Abstract

Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surprisingly, we prove that even simple, individually stable models can generate relatively unstable comparisons, calling into question the validity of CV inference. Specifically, we show that the Lasso and its close cousin, soft-thresholding, generate relatively unstable comparisons and invalid CV inferences, even in the most favorable of learning settings and when both models are individually stable. These findings highlight the importance of verifying relative stability before deploying CV for model comparison.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bayle26a, title = {The Relative Instability of Model Comparison with Cross-validation}, author = {Bayle, Alexandre and Janson, Lucas and Mackey, Lester}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {7144--7178}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bayle26a/bayle26a.pdf}, url = {https://proceedings.mlr.press/v306/bayle26a.html}, abstract = {Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surprisingly, we prove that even simple, individually stable models can generate relatively unstable comparisons, calling into question the validity of CV inference. Specifically, we show that the Lasso and its close cousin, soft-thresholding, generate relatively unstable comparisons and invalid CV inferences, even in the most favorable of learning settings and when both models are individually stable. These findings highlight the importance of verifying relative stability before deploying CV for model comparison.} }
Endnote
%0 Conference Paper %T The Relative Instability of Model Comparison with Cross-validation %A Alexandre Bayle %A Lucas Janson %A Lester Mackey %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bayle26a %I PMLR %P 7144--7178 %U https://proceedings.mlr.press/v306/bayle26a.html %V 306 %X Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surprisingly, we prove that even simple, individually stable models can generate relatively unstable comparisons, calling into question the validity of CV inference. Specifically, we show that the Lasso and its close cousin, soft-thresholding, generate relatively unstable comparisons and invalid CV inferences, even in the most favorable of learning settings and when both models are individually stable. These findings highlight the importance of verifying relative stability before deploying CV for model comparison.
APA
Bayle, A., Janson, L. & Mackey, L.. (2026). The Relative Instability of Model Comparison with Cross-validation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:7144-7178 Available from https://proceedings.mlr.press/v306/bayle26a.html.

Related Material