Function-Space Learning Rates

Edward Milsom, Ben Anson, Laurence Aitchison
Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:44225-44251, 2025.

Abstract

We consider layerwise function-space learning rates, which measure the magnitude of the change in a neural network’s output function in response to an update to a parameter tensor. This contrasts with traditional learning rates, which describe the magnitude of changes in parameter space. We develop efficient methods to measure and set function-space learning rates in arbitrary neural networks, requiring only minimal computational overhead through a few additional backward passes that can be performed at the start of, or periodically during, training. We demonstrate two key applications: (1) analysing the dynamics of standard neural network optimisers in function space, rather than parameter space, and (2) introducing FLeRM (Function-space Learning Rate Matching), a novel approach to hyperparameter transfer across model scales. FLeRM records function-space learning rates while training a small, cheap base model, then automatically adjusts parameter-space layerwise learning rates when training larger models to maintain consistent function-space updates. FLeRM gives hyperparameter transfer across model width, depth, initialisation scale, and LoRA rank in various architectures including MLPs with residual connections and transformers with different layer normalisation schemes.

Cite this Paper


BibTeX
@InProceedings{pmlr-v267-milsom25a, title = {Function-Space Learning Rates}, author = {Milsom, Edward and Anson, Ben and Aitchison, Laurence}, booktitle = {Proceedings of the 42nd International Conference on Machine Learning}, pages = {44225--44251}, year = {2025}, editor = {Singh, Aarti and Fazel, Maryam and Hsu, Daniel and Lacoste-Julien, Simon and Berkenkamp, Felix and Maharaj, Tegan and Wagstaff, Kiri and Zhu, Jerry}, volume = {267}, series = {Proceedings of Machine Learning Research}, month = {13--19 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v267/main/assets/milsom25a/milsom25a.pdf}, url = {https://proceedings.mlr.press/v267/milsom25a.html}, abstract = {We consider layerwise function-space learning rates, which measure the magnitude of the change in a neural network’s output function in response to an update to a parameter tensor. This contrasts with traditional learning rates, which describe the magnitude of changes in parameter space. We develop efficient methods to measure and set function-space learning rates in arbitrary neural networks, requiring only minimal computational overhead through a few additional backward passes that can be performed at the start of, or periodically during, training. We demonstrate two key applications: (1) analysing the dynamics of standard neural network optimisers in function space, rather than parameter space, and (2) introducing FLeRM (Function-space Learning Rate Matching), a novel approach to hyperparameter transfer across model scales. FLeRM records function-space learning rates while training a small, cheap base model, then automatically adjusts parameter-space layerwise learning rates when training larger models to maintain consistent function-space updates. FLeRM gives hyperparameter transfer across model width, depth, initialisation scale, and LoRA rank in various architectures including MLPs with residual connections and transformers with different layer normalisation schemes.} }
Endnote
%0 Conference Paper %T Function-Space Learning Rates %A Edward Milsom %A Ben Anson %A Laurence Aitchison %B Proceedings of the 42nd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2025 %E Aarti Singh %E Maryam Fazel %E Daniel Hsu %E Simon Lacoste-Julien %E Felix Berkenkamp %E Tegan Maharaj %E Kiri Wagstaff %E Jerry Zhu %F pmlr-v267-milsom25a %I PMLR %P 44225--44251 %U https://proceedings.mlr.press/v267/milsom25a.html %V 267 %X We consider layerwise function-space learning rates, which measure the magnitude of the change in a neural network’s output function in response to an update to a parameter tensor. This contrasts with traditional learning rates, which describe the magnitude of changes in parameter space. We develop efficient methods to measure and set function-space learning rates in arbitrary neural networks, requiring only minimal computational overhead through a few additional backward passes that can be performed at the start of, or periodically during, training. We demonstrate two key applications: (1) analysing the dynamics of standard neural network optimisers in function space, rather than parameter space, and (2) introducing FLeRM (Function-space Learning Rate Matching), a novel approach to hyperparameter transfer across model scales. FLeRM records function-space learning rates while training a small, cheap base model, then automatically adjusts parameter-space layerwise learning rates when training larger models to maintain consistent function-space updates. FLeRM gives hyperparameter transfer across model width, depth, initialisation scale, and LoRA rank in various architectures including MLPs with residual connections and transformers with different layer normalisation schemes.
APA
Milsom, E., Anson, B. & Aitchison, L.. (2025). Function-Space Learning Rates. Proceedings of the 42nd International Conference on Machine Learning, in Proceedings of Machine Learning Research 267:44225-44251 Available from https://proceedings.mlr.press/v267/milsom25a.html.

Related Material