Improved Stochastic Optimization of LogSumExp

Egor Gladin, Alexey Kroshnin, Jia-Jie Zhu, Pavel Dvurechensky
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:35176-35196, 2026.

Abstract

The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In practice, when the number of exponential terms inside the logarithm is large or infinite, optimization becomes challenging since computing the gradient requires differentiating every term. We propose a novel convexity- and smoothness-preserving approximation to LogSumExp that can be efficiently optimized using stochastic gradient methods. This approximation is rooted in a sound modification of the KL divergence in the dual, resulting in a new $f$-divergence called the Safe KL divergence. Our experiments and theoretical analysis of the LogSumExp-based stochastic optimization, arising in DRO and continuous OT, demonstrate the advantages of our approach over existing baselines.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gladin26a, title = {Improved Stochastic Optimization of {L}og{S}um{E}xp}, author = {Gladin, Egor and Kroshnin, Alexey and Zhu, Jia-Jie and Dvurechensky, Pavel}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {35176--35196}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gladin26a/gladin26a.pdf}, url = {https://proceedings.mlr.press/v306/gladin26a.html}, abstract = {The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In practice, when the number of exponential terms inside the logarithm is large or infinite, optimization becomes challenging since computing the gradient requires differentiating every term. We propose a novel convexity- and smoothness-preserving approximation to LogSumExp that can be efficiently optimized using stochastic gradient methods. This approximation is rooted in a sound modification of the KL divergence in the dual, resulting in a new $f$-divergence called the Safe KL divergence. Our experiments and theoretical analysis of the LogSumExp-based stochastic optimization, arising in DRO and continuous OT, demonstrate the advantages of our approach over existing baselines.} }
Endnote
%0 Conference Paper %T Improved Stochastic Optimization of LogSumExp %A Egor Gladin %A Alexey Kroshnin %A Jia-Jie Zhu %A Pavel Dvurechensky %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gladin26a %I PMLR %P 35176--35196 %U https://proceedings.mlr.press/v306/gladin26a.html %V 306 %X The LogSumExp function, dual to the Kullback-Leibler (KL) divergence, plays a central role in many important optimization problems, including entropy-regularized optimal transport (OT) and distributionally robust optimization (DRO). In practice, when the number of exponential terms inside the logarithm is large or infinite, optimization becomes challenging since computing the gradient requires differentiating every term. We propose a novel convexity- and smoothness-preserving approximation to LogSumExp that can be efficiently optimized using stochastic gradient methods. This approximation is rooted in a sound modification of the KL divergence in the dual, resulting in a new $f$-divergence called the Safe KL divergence. Our experiments and theoretical analysis of the LogSumExp-based stochastic optimization, arising in DRO and continuous OT, demonstrate the advantages of our approach over existing baselines.
APA
Gladin, E., Kroshnin, A., Zhu, J. & Dvurechensky, P.. (2026). Improved Stochastic Optimization of LogSumExp. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:35176-35196 Available from https://proceedings.mlr.press/v306/gladin26a.html.

Related Material