On the Finite-Sample Bias of Minimizing Expected Wasserstein Loss Between Empirical Distributions

Cheongjae Jang, Yung-Kyun Noh
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4420-4428, 2026.

Abstract

We show that minimizing the expected Wasserstein loss between empirical distributions can lead to biased parameter estimates in the finite-sample regime. Remarkably, such bias arises even in well-specified settings where both empirical distributions are drawn from the same parametric family: unlike maximum likelihood estimation—understood here as maximizing the expected log-likelihood—optimizing one parameter while fixing another fails to recover the true fixed value. We derive closed-form expressions for the expected Wasserstein loss in one dimension and, focusing on location–scale models, provide an analytic characterization of the bias. This analysis reveals that finite-sample bias occurs whenever the expected loss varies along the diagonal subspace where parameter values coincide, and we propose a simple correction scheme that removes this effect. We extend our analysis to misspecified models and the Sinkhorn divergence, demonstrating that finite-sample bias persists in more practical settings. Experiments on synthetic and real data confirm that stochastic optimization of Wasserstein-based objectives converges to biased solutions, and validate the effectiveness of the proposed correction scheme.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-jang26b, title = { On the Finite-Sample Bias of Minimizing Expected Wasserstein Loss Between Empirical Distributions }, author = {Jang, Cheongjae and Noh, Yung-Kyun}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4420--4428}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/jang26b/jang26b.pdf}, url = {https://proceedings.mlr.press/v300/jang26b.html}, abstract = { We show that minimizing the expected Wasserstein loss between empirical distributions can lead to biased parameter estimates in the finite-sample regime. Remarkably, such bias arises even in well-specified settings where both empirical distributions are drawn from the same parametric family: unlike maximum likelihood estimation—understood here as maximizing the expected log-likelihood—optimizing one parameter while fixing another fails to recover the true fixed value. We derive closed-form expressions for the expected Wasserstein loss in one dimension and, focusing on location–scale models, provide an analytic characterization of the bias. This analysis reveals that finite-sample bias occurs whenever the expected loss varies along the diagonal subspace where parameter values coincide, and we propose a simple correction scheme that removes this effect. We extend our analysis to misspecified models and the Sinkhorn divergence, demonstrating that finite-sample bias persists in more practical settings. Experiments on synthetic and real data confirm that stochastic optimization of Wasserstein-based objectives converges to biased solutions, and validate the effectiveness of the proposed correction scheme. } }
Endnote
%0 Conference Paper %T On the Finite-Sample Bias of Minimizing Expected Wasserstein Loss Between Empirical Distributions %A Cheongjae Jang %A Yung-Kyun Noh %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-jang26b %I PMLR %P 4420--4428 %U https://proceedings.mlr.press/v300/jang26b.html %V 300 %X We show that minimizing the expected Wasserstein loss between empirical distributions can lead to biased parameter estimates in the finite-sample regime. Remarkably, such bias arises even in well-specified settings where both empirical distributions are drawn from the same parametric family: unlike maximum likelihood estimation—understood here as maximizing the expected log-likelihood—optimizing one parameter while fixing another fails to recover the true fixed value. We derive closed-form expressions for the expected Wasserstein loss in one dimension and, focusing on location–scale models, provide an analytic characterization of the bias. This analysis reveals that finite-sample bias occurs whenever the expected loss varies along the diagonal subspace where parameter values coincide, and we propose a simple correction scheme that removes this effect. We extend our analysis to misspecified models and the Sinkhorn divergence, demonstrating that finite-sample bias persists in more practical settings. Experiments on synthetic and real data confirm that stochastic optimization of Wasserstein-based objectives converges to biased solutions, and validate the effectiveness of the proposed correction scheme.
APA
Jang, C. & Noh, Y.. (2026). On the Finite-Sample Bias of Minimizing Expected Wasserstein Loss Between Empirical Distributions . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4420-4428 Available from https://proceedings.mlr.press/v300/jang26b.html.

Related Material