On Global Convergence Rates for Federated Softmax Policy Gradient Under HeterogeneousEnvironments

Safwan Labbi, Paul Mangold, Daniil Tiapkin, Eric Moulines
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4960-4968, 2026.

Abstract

We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient ($\texttt{FedPG}$) with local training. We show that $\texttt{FedPG}$ converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient \emph{with explicit constants}, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the {Ł}ojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-labbi26a, title = { On Global Convergence Rates for Federated Softmax Policy Gradient Under HeterogeneousEnvironments }, author = {Labbi, Safwan and Mangold, Paul and Tiapkin, Daniil and Moulines, Eric}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4960--4968}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/labbi26a/labbi26a.pdf}, url = {https://proceedings.mlr.press/v300/labbi26a.html}, abstract = { We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient ($\texttt{FedPG}$) with local training. We show that $\texttt{FedPG}$ converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient \emph{with explicit constants}, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the {Ł}ojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies. } }
Endnote
%0 Conference Paper %T On Global Convergence Rates for Federated Softmax Policy Gradient Under HeterogeneousEnvironments %A Safwan Labbi %A Paul Mangold %A Daniil Tiapkin %A Eric Moulines %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-labbi26a %I PMLR %P 4960--4968 %U https://proceedings.mlr.press/v300/labbi26a.html %V 300 %X We provide global convergence rates for vanilla and entropy-regularized federated softmax stochastic policy gradient ($\texttt{FedPG}$) with local training. We show that $\texttt{FedPG}$ converges to a near-optimal policy in terms of the average agent value, with a gap controlled by the level of heterogeneity. Remarkably, we obtain the first convergence rates for entropy-regularized policy gradient \emph{with explicit constants}, leveraging a projection-like operator. Our results build upon a new analysis of federated averaging for non-convex objectives, based on the observation that the {Ł}ojasiewicz-type inequalities from the single-agent setting (Mei et al., 2020) do not hold for the federated objective. This uncovers a fundamental difference between single-agent and federated reinforcement learning: while single-agent optimal policies can be deterministic, federated objectives may inherently require stochastic policies.
APA
Labbi, S., Mangold, P., Tiapkin, D. & Moulines, E.. (2026). On Global Convergence Rates for Federated Softmax Policy Gradient Under HeterogeneousEnvironments . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4960-4968 Available from https://proceedings.mlr.press/v300/labbi26a.html.

Related Material