Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning

Tolga Dimlioglu, Anna Ewa Choromanska
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1378-1422, 2026.

Abstract

We study centralized distributed data parallel training of deep neural networks (DNNs), aiming to improve the trade-off between communication efficiency and model performance of local gradient methods. Motivated by the flat-minima hypothesis, we first introduce a simple sharpness measure, Inverse Mean Valley, and show it strongly correlates with the generalization gap of DNNs. We then incorporate an efficient relaxation of this measure into the distributed objective as a lightweight regularizer that encourages workers to seek wide minima collaboratively. The regularizer exerts a pushing force that counteracts the consensus step pulling the workers together, giving rise to the Distributed Pull-Push Force ({DPPF}) algorithm. Empirically, {DPPF} generalizes better than other local gradient methods and synchronous gradient averaging while maintaining communication efficiency. In addition, our loss landscape visualizations confirm the ability of {DPPF} to locate flatter minima. Theoretically, we show that {DPPF} drives workers to span flat valleys with valley width governed by push–pull strengths, it yields self-stabilizing dynamics, it obeys generalization guarantees that depend on valley width, and it converges in the non-convex setting.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-dimlioglu26a, title = {Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning}, author = {Dimlioglu, Tolga and Choromanska, Anna Ewa}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {1378--1422}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/dimlioglu26a/dimlioglu26a.pdf}, url = {https://proceedings.mlr.press/v337/dimlioglu26a.html}, abstract = {We study centralized distributed data parallel training of deep neural networks (DNNs), aiming to improve the trade-off between communication efficiency and model performance of local gradient methods. Motivated by the flat-minima hypothesis, we first introduce a simple sharpness measure, Inverse Mean Valley, and show it strongly correlates with the generalization gap of DNNs. We then incorporate an efficient relaxation of this measure into the distributed objective as a lightweight regularizer that encourages workers to seek wide minima collaboratively. The regularizer exerts a pushing force that counteracts the consensus step pulling the workers together, giving rise to the Distributed Pull-Push Force ({DPPF}) algorithm. Empirically, {DPPF} generalizes better than other local gradient methods and synchronous gradient averaging while maintaining communication efficiency. In addition, our loss landscape visualizations confirm the ability of {DPPF} to locate flatter minima. Theoretically, we show that {DPPF} drives workers to span flat valleys with valley width governed by push–pull strengths, it yields self-stabilizing dynamics, it obeys generalization guarantees that depend on valley width, and it converges in the non-convex setting.} }
Endnote
%0 Conference Paper %T Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning %A Tolga Dimlioglu %A Anna Ewa Choromanska %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-dimlioglu26a %I PMLR %P 1378--1422 %U https://proceedings.mlr.press/v337/dimlioglu26a.html %V 337 %X We study centralized distributed data parallel training of deep neural networks (DNNs), aiming to improve the trade-off between communication efficiency and model performance of local gradient methods. Motivated by the flat-minima hypothesis, we first introduce a simple sharpness measure, Inverse Mean Valley, and show it strongly correlates with the generalization gap of DNNs. We then incorporate an efficient relaxation of this measure into the distributed objective as a lightweight regularizer that encourages workers to seek wide minima collaboratively. The regularizer exerts a pushing force that counteracts the consensus step pulling the workers together, giving rise to the Distributed Pull-Push Force ({DPPF}) algorithm. Empirically, {DPPF} generalizes better than other local gradient methods and synchronous gradient averaging while maintaining communication efficiency. In addition, our loss landscape visualizations confirm the ability of {DPPF} to locate flatter minima. Theoretically, we show that {DPPF} drives workers to span flat valleys with valley width governed by push–pull strengths, it yields self-stabilizing dynamics, it obeys generalization guarantees that depend on valley width, and it converges in the non-convex setting.
APA
Dimlioglu, T. & Choromanska, A.E.. (2026). Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:1378-1422 Available from https://proceedings.mlr.press/v337/dimlioglu26a.html.

Related Material