Shift is Good: Mismatched Data Mixing Improves Test Performance

Marko Medvedev, Kaifeng Lyu, Zhiyuan Li, Nathan Srebro
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4690-4698, 2026.

Abstract

We consider training and testing on mixture distributions with different training and test proportions. We show that in many settings, and in some sense generically, distribution shift can be beneficial, and test performance can improve due to mismatched training proportions, even if the components are unrelated and with no transfer between components. In a variety of scenarios, we identify the optimal training proportions and the extent to which such distribution shift can be beneficial. We show how the same analysis applies also to a compositional setting with differing distribution of component “skills” at training and test.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-medvedev26a, title = { Shift is Good: Mismatched Data Mixing Improves Test Performance }, author = {Medvedev, Marko and Lyu, Kaifeng and Li, Zhiyuan and Srebro, Nathan}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4690--4698}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/medvedev26a/medvedev26a.pdf}, url = {https://proceedings.mlr.press/v300/medvedev26a.html}, abstract = { We consider training and testing on mixture distributions with different training and test proportions. We show that in many settings, and in some sense generically, distribution shift can be beneficial, and test performance can improve due to mismatched training proportions, even if the components are unrelated and with no transfer between components. In a variety of scenarios, we identify the optimal training proportions and the extent to which such distribution shift can be beneficial. We show how the same analysis applies also to a compositional setting with differing distribution of component “skills” at training and test. } }
Endnote
%0 Conference Paper %T Shift is Good: Mismatched Data Mixing Improves Test Performance %A Marko Medvedev %A Kaifeng Lyu %A Zhiyuan Li %A Nathan Srebro %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-medvedev26a %I PMLR %P 4690--4698 %U https://proceedings.mlr.press/v300/medvedev26a.html %V 300 %X We consider training and testing on mixture distributions with different training and test proportions. We show that in many settings, and in some sense generically, distribution shift can be beneficial, and test performance can improve due to mismatched training proportions, even if the components are unrelated and with no transfer between components. In a variety of scenarios, we identify the optimal training proportions and the extent to which such distribution shift can be beneficial. We show how the same analysis applies also to a compositional setting with differing distribution of component “skills” at training and test.
APA
Medvedev, M., Lyu, K., Li, Z. & Srebro, N.. (2026). Shift is Good: Mismatched Data Mixing Improves Test Performance . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4690-4698 Available from https://proceedings.mlr.press/v300/medvedev26a.html.

Related Material