Beyond Pooling: Matching for Robust Generalization Under Data Heterogeneity

Ayush Roy, Rudrasis Chakraborty, Lav R. Varshney, Vishnu Suresh Lokhande
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:82-90, 2026.

Abstract

Pooling heterogeneous datasets across domains is a common strategy in representation learning, but naive pooling can amplify distributional asymmetries and yield biased estimators, especially in settings where zero-shot generalization is required. We propose a matching framework that selects samples relative to an adaptive centroid and iteratively refines the representation distribution. The double robustness and the propensity score matching for the inclusion of data domains make matching more robust than naive pooling and uniform subsampling by filtering out the confounding domains (the main cause of heterogeneity). Theoretical and empirical analyses show that, unlike naive pooling or uniform subsampling, matching achieves better results under asymmetric meta-distributions, which are also extended to non-Gaussian and multimodal real-world settings. Most importantly, we show that these improvements translate to zero-shot medical anomaly detection, one of the extreme forms of data heterogeneity and asymmetry. The code is available on Github.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-roy26a, title = { Beyond Pooling: Matching for Robust Generalization Under Data Heterogeneity }, author = {Roy, Ayush and Chakraborty, Rudrasis and Varshney, Lav R. and Lokhande, Vishnu Suresh}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {82--90}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/roy26a/roy26a.pdf}, url = {https://proceedings.mlr.press/v300/roy26a.html}, abstract = { Pooling heterogeneous datasets across domains is a common strategy in representation learning, but naive pooling can amplify distributional asymmetries and yield biased estimators, especially in settings where zero-shot generalization is required. We propose a matching framework that selects samples relative to an adaptive centroid and iteratively refines the representation distribution. The double robustness and the propensity score matching for the inclusion of data domains make matching more robust than naive pooling and uniform subsampling by filtering out the confounding domains (the main cause of heterogeneity). Theoretical and empirical analyses show that, unlike naive pooling or uniform subsampling, matching achieves better results under asymmetric meta-distributions, which are also extended to non-Gaussian and multimodal real-world settings. Most importantly, we show that these improvements translate to zero-shot medical anomaly detection, one of the extreme forms of data heterogeneity and asymmetry. The code is available on Github. } }
Endnote
%0 Conference Paper %T Beyond Pooling: Matching for Robust Generalization Under Data Heterogeneity %A Ayush Roy %A Rudrasis Chakraborty %A Lav R. Varshney %A Vishnu Suresh Lokhande %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-roy26a %I PMLR %P 82--90 %U https://proceedings.mlr.press/v300/roy26a.html %V 300 %X Pooling heterogeneous datasets across domains is a common strategy in representation learning, but naive pooling can amplify distributional asymmetries and yield biased estimators, especially in settings where zero-shot generalization is required. We propose a matching framework that selects samples relative to an adaptive centroid and iteratively refines the representation distribution. The double robustness and the propensity score matching for the inclusion of data domains make matching more robust than naive pooling and uniform subsampling by filtering out the confounding domains (the main cause of heterogeneity). Theoretical and empirical analyses show that, unlike naive pooling or uniform subsampling, matching achieves better results under asymmetric meta-distributions, which are also extended to non-Gaussian and multimodal real-world settings. Most importantly, we show that these improvements translate to zero-shot medical anomaly detection, one of the extreme forms of data heterogeneity and asymmetry. The code is available on Github.
APA
Roy, A., Chakraborty, R., Varshney, L.R. & Lokhande, V.S.. (2026). Beyond Pooling: Matching for Robust Generalization Under Data Heterogeneity . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:82-90 Available from https://proceedings.mlr.press/v300/roy26a.html.

Related Material