Bregman contrastive estimation as a general framework for learning unnormalized models: Efficiency, robustness and optimal noise

Yuto Fujii, Hiroaki Sasaki, Takafumi Kanamori
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1604-1627, 2026.

Abstract

Learning unnormalized models is an ubiquitous problem in modern statistics and machine learning. Noise contrastive estimation (NCE) takes a popular approach based on solving a binary classification problem: Unnormalized models are learned by discriminating input data with artificially generated noise data. This paper follows this approach, but proposes a general framework based on the {Bregman} divergence. By selecting the convex function in the divergence, our framework includes existing methods as special cases, and novel variants of NCE can be also derived. Learning unnormalized models in the proposed framework is performed by estimating the posterior probability in the binary classification, i.e., the composition of the logistic function and the ratio of data and noise distributions. Due to the boundedness of the logistic function, our framework is advantageous in outlier-robustness. In fact, we theoretically prove that robust estimation is possible in our framework under a variety of convex functions in the {Bregman} divergence. Furthermore, the asymptotic efficiency is also investigated, implying a trade-off between efficiency and robustness in our framework. Inspired by these results, we derive the optimal noise distribution under some constraints for robustness. Finally, we numerically demonstrate the robustness of the novel variants of NCE and the derived optimal noise distribution.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-fujii26a, title = {{Bregman} contrastive estimation as a general framework for learning unnormalized models: Efficiency, robustness and optimal noise}, author = {Fujii, Yuto and Sasaki, Hiroaki and Kanamori, Takafumi}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {1604--1627}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/fujii26a/fujii26a.pdf}, url = {https://proceedings.mlr.press/v337/fujii26a.html}, abstract = {Learning unnormalized models is an ubiquitous problem in modern statistics and machine learning. Noise contrastive estimation (NCE) takes a popular approach based on solving a binary classification problem: Unnormalized models are learned by discriminating input data with artificially generated noise data. This paper follows this approach, but proposes a general framework based on the {Bregman} divergence. By selecting the convex function in the divergence, our framework includes existing methods as special cases, and novel variants of NCE can be also derived. Learning unnormalized models in the proposed framework is performed by estimating the posterior probability in the binary classification, i.e., the composition of the logistic function and the ratio of data and noise distributions. Due to the boundedness of the logistic function, our framework is advantageous in outlier-robustness. In fact, we theoretically prove that robust estimation is possible in our framework under a variety of convex functions in the {Bregman} divergence. Furthermore, the asymptotic efficiency is also investigated, implying a trade-off between efficiency and robustness in our framework. Inspired by these results, we derive the optimal noise distribution under some constraints for robustness. Finally, we numerically demonstrate the robustness of the novel variants of NCE and the derived optimal noise distribution.} }
Endnote
%0 Conference Paper %T Bregman contrastive estimation as a general framework for learning unnormalized models: Efficiency, robustness and optimal noise %A Yuto Fujii %A Hiroaki Sasaki %A Takafumi Kanamori %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-fujii26a %I PMLR %P 1604--1627 %U https://proceedings.mlr.press/v337/fujii26a.html %V 337 %X Learning unnormalized models is an ubiquitous problem in modern statistics and machine learning. Noise contrastive estimation (NCE) takes a popular approach based on solving a binary classification problem: Unnormalized models are learned by discriminating input data with artificially generated noise data. This paper follows this approach, but proposes a general framework based on the {Bregman} divergence. By selecting the convex function in the divergence, our framework includes existing methods as special cases, and novel variants of NCE can be also derived. Learning unnormalized models in the proposed framework is performed by estimating the posterior probability in the binary classification, i.e., the composition of the logistic function and the ratio of data and noise distributions. Due to the boundedness of the logistic function, our framework is advantageous in outlier-robustness. In fact, we theoretically prove that robust estimation is possible in our framework under a variety of convex functions in the {Bregman} divergence. Furthermore, the asymptotic efficiency is also investigated, implying a trade-off between efficiency and robustness in our framework. Inspired by these results, we derive the optimal noise distribution under some constraints for robustness. Finally, we numerically demonstrate the robustness of the novel variants of NCE and the derived optimal noise distribution.
APA
Fujii, Y., Sasaki, H. & Kanamori, T.. (2026). Bregman contrastive estimation as a general framework for learning unnormalized models: Efficiency, robustness and optimal noise. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:1604-1627 Available from https://proceedings.mlr.press/v337/fujii26a.html.

Related Material