A Semi-Supervised Kernel Two-Sample Test

Gyumin Lee, Shubhanshu Shekhar, Ilmun Kim
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:838-846, 2026.

Abstract

We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However, incorporating covariates potentially breaks the exchangeability assumption under the null, which further complicates a calibration procedure. To address these issues, we propose a semi-supervised method that produces a test statistic with asymptotic normality, while effectively integrating additional information from covariates. Our test is straightforward to calibrate due to the asymptotic normality under the null and achieves asymptotic power that is often much higher than existing kernel tests without covariates. Furthermore, we formally show that the proposed method is consistent in power against fixed and local alternatives. Simulations confirm the practical and theoretical strengths of our approach.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-lee26b, title = { A Semi-Supervised Kernel Two-Sample Test }, author = {Lee, Gyumin and Shekhar, Shubhanshu and Kim, Ilmun}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {838--846}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/lee26b/lee26b.pdf}, url = {https://proceedings.mlr.press/v300/lee26b.html}, abstract = { We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However, incorporating covariates potentially breaks the exchangeability assumption under the null, which further complicates a calibration procedure. To address these issues, we propose a semi-supervised method that produces a test statistic with asymptotic normality, while effectively integrating additional information from covariates. Our test is straightforward to calibrate due to the asymptotic normality under the null and achieves asymptotic power that is often much higher than existing kernel tests without covariates. Furthermore, we formally show that the proposed method is consistent in power against fixed and local alternatives. Simulations confirm the practical and theoretical strengths of our approach. } }
Endnote
%0 Conference Paper %T A Semi-Supervised Kernel Two-Sample Test %A Gyumin Lee %A Shubhanshu Shekhar %A Ilmun Kim %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-lee26b %I PMLR %P 838--846 %U https://proceedings.mlr.press/v300/lee26b.html %V 300 %X We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However, incorporating covariates potentially breaks the exchangeability assumption under the null, which further complicates a calibration procedure. To address these issues, we propose a semi-supervised method that produces a test statistic with asymptotic normality, while effectively integrating additional information from covariates. Our test is straightforward to calibrate due to the asymptotic normality under the null and achieves asymptotic power that is often much higher than existing kernel tests without covariates. Furthermore, we formally show that the proposed method is consistent in power against fixed and local alternatives. Simulations confirm the practical and theoretical strengths of our approach.
APA
Lee, G., Shekhar, S. & Kim, I.. (2026). A Semi-Supervised Kernel Two-Sample Test . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:838-846 Available from https://proceedings.mlr.press/v300/lee26b.html.

Related Material