Learning rate collapse prevents training recurrent neural networks at scale

Bariscan Kurtkaya, Mehmet Harmanli, Alperen Cimen, Andy Alexander, Nina Miolane, Fatih Dinc, Yucel Yemez
Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations, PMLR 282:282-297, 2026.

Abstract

Recurrent neural networks (RNNs) are central to modeling neural computation in systems neuroscience, yet the principles that enable their stable and efficient training at large scales remain poorly understood. Seminal work in machine learning predicts that the effective learning rate should shrink with the size of feedforward networks. Here, we demonstrate an analogous phenomenon, termed learning rate collapse, in which the maximum trainable learning rate decreases inversely with the number of neurons. This behavior can be mitigated partially by scaling parameters with the inverse of network, though learning still takes longer for larger networks. These limits are further compounded by severe memory demands, which together make training large RNNs both unstable and computationally costly. As a proof of principle for mitigating learning rate collapse, we study the learning process of low-rank networks, which enforces a low-dimensional geometry in RNN representations. These results situate learning rate collapse within a broader lineage of scaling analyses in RNNs, with potential solutions likely to come from future work that incorporates careful consideration of symmetry and geometry in neural representations.

Cite this Paper


BibTeX
@InProceedings{pmlr-v282-kurtkaya26a, title = {Learning rate collapse prevents training recurrent neural networks at scale}, author = {Kurtkaya, Bariscan and Harmanli, Mehmet and Cimen, Alperen and Alexander, Andy and Miolane, Nina and Dinc, Fatih and Yemez, Yucel}, booktitle = {Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations}, pages = {282--297}, year = {2026}, editor = {Acosta, Francisco and Azeglio, Simone and Tolooshams, Bahareh and van de Geijn, Chase and Shewmake, Christian and Sanborn, Sophia and Miolane, Nina}, volume = {282}, series = {Proceedings of Machine Learning Research}, month = {14 Dec 2024--07 Dec 2025}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v282/main/assets/kurtkaya26a/kurtkaya26a.pdf}, url = {https://proceedings.mlr.press/v282/kurtkaya26a.html}, abstract = {Recurrent neural networks (RNNs) are central to modeling neural computation in systems neuroscience, yet the principles that enable their stable and efficient training at large scales remain poorly understood. Seminal work in machine learning predicts that the effective learning rate should shrink with the size of feedforward networks. Here, we demonstrate an analogous phenomenon, termed learning rate collapse, in which the maximum trainable learning rate decreases inversely with the number of neurons. This behavior can be mitigated partially by scaling parameters with the inverse of network, though learning still takes longer for larger networks. These limits are further compounded by severe memory demands, which together make training large RNNs both unstable and computationally costly. As a proof of principle for mitigating learning rate collapse, we study the learning process of low-rank networks, which enforces a low-dimensional geometry in RNN representations. These results situate learning rate collapse within a broader lineage of scaling analyses in RNNs, with potential solutions likely to come from future work that incorporates careful consideration of symmetry and geometry in neural representations.} }
Endnote
%0 Conference Paper %T Learning rate collapse prevents training recurrent neural networks at scale %A Bariscan Kurtkaya %A Mehmet Harmanli %A Alperen Cimen %A Andy Alexander %A Nina Miolane %A Fatih Dinc %A Yucel Yemez %B Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations %C Proceedings of Machine Learning Research %D 2026 %E Francisco Acosta %E Simone Azeglio %E Bahareh Tolooshams %E Chase van de Geijn %E Christian Shewmake %E Sophia Sanborn %E Nina Miolane %F pmlr-v282-kurtkaya26a %I PMLR %P 282--297 %U https://proceedings.mlr.press/v282/kurtkaya26a.html %V 282 %X Recurrent neural networks (RNNs) are central to modeling neural computation in systems neuroscience, yet the principles that enable their stable and efficient training at large scales remain poorly understood. Seminal work in machine learning predicts that the effective learning rate should shrink with the size of feedforward networks. Here, we demonstrate an analogous phenomenon, termed learning rate collapse, in which the maximum trainable learning rate decreases inversely with the number of neurons. This behavior can be mitigated partially by scaling parameters with the inverse of network, though learning still takes longer for larger networks. These limits are further compounded by severe memory demands, which together make training large RNNs both unstable and computationally costly. As a proof of principle for mitigating learning rate collapse, we study the learning process of low-rank networks, which enforces a low-dimensional geometry in RNN representations. These results situate learning rate collapse within a broader lineage of scaling analyses in RNNs, with potential solutions likely to come from future work that incorporates careful consideration of symmetry and geometry in neural representations.
APA
Kurtkaya, B., Harmanli, M., Cimen, A., Alexander, A., Miolane, N., Dinc, F. & Yemez, Y.. (2026). Learning rate collapse prevents training recurrent neural networks at scale. Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations, in Proceedings of Machine Learning Research 282:282-297 Available from https://proceedings.mlr.press/v282/kurtkaya26a.html.

Related Material