The Affine Divergence: Aligning Activation Updates Beyond Normalisation

George Bird
Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling, PMLR 326:45-74, 2026.

Abstract

A systematic mismatch exists between mathematically ideal and effective activation updates during gradient descent. As intended, parameters update in their direction of steepest descent. However, activations are argued to constitute a more directly impactful quantity to prioritise in optimisation, as they are closer to the loss in the computational graph and carry sample-dependent information through the network. Yet their propagated updates do not take the optimal steepest-descent step. These quantities exhibit non-ideal sample-wise scaling across affine, convolutional, and attention layers. Solutions to correct for this are trivial and, incidentally, derive normalisation from first principles despite motivational independence. Consequently, such considerations offer a fresh, conceptual reframe of normalisation’s action, with auxiliary experiments bolstering this mechanistic interpretation. Moreover, this analysis makes clear a second possibility: a solution that is functionally distinct from modern normalisations, without scale invariance, yet remains empirically successful — an alternative to the affine map. This outperforms conventional normalisers across several tests. This generalises to convolution via a new functional form, “PatchNorm”, a compositionally inseparable normaliser. Together, these provide an alternative mechanistic framework that both adds to and counters some of the discussion of normalisation. Further, it is argued that normalisers are better decomposed into activation-function-like maps with parameterised scaling. Overall, this constitutes a theoretically principled approach that yields new functions with empirical validation and raises questions about the affine + nonlinear approach.

Cite this Paper


BibTeX
@InProceedings{pmlr-v326-bird26a, title = {{T}he {A}ffine {D}ivergence: {A}ligning {A}ctivation {U}pdates {B}eyond {N}ormalisation}, author = {Bird, George}, booktitle = {Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling}, pages = {45--74}, year = {2026}, editor = {Pouplin, Alison and Vadgama, Sharvaree and Bekkers, Erik and Kaba, Sékou-Oumar and Lawrence, Hannah and Lecha, Manuel and Baker, Elizabeth and Suk, Julian and Walters, Robin and Tomczak, Jakub and Jegelka, Stefanie}, volume = {326}, series = {Proceedings of Machine Learning Research}, month = {26 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v326/main/assets/bird26a/bird26a.pdf}, url = {https://proceedings.mlr.press/v326/bird26a.html}, abstract = {A systematic mismatch exists between mathematically ideal and effective activation updates during gradient descent. As intended, parameters update in their direction of steepest descent. However, activations are argued to constitute a more directly impactful quantity to prioritise in optimisation, as they are closer to the loss in the computational graph and carry sample-dependent information through the network. Yet their propagated updates do not take the optimal steepest-descent step. These quantities exhibit non-ideal sample-wise scaling across affine, convolutional, and attention layers. Solutions to correct for this are trivial and, incidentally, derive normalisation from first principles despite motivational independence. Consequently, such considerations offer a fresh, conceptual reframe of normalisation’s action, with auxiliary experiments bolstering this mechanistic interpretation. Moreover, this analysis makes clear a second possibility: a solution that is functionally distinct from modern normalisations, without scale invariance, yet remains empirically successful — an alternative to the affine map. This outperforms conventional normalisers across several tests. This generalises to convolution via a new functional form, “PatchNorm”, a compositionally inseparable normaliser. Together, these provide an alternative mechanistic framework that both adds to and counters some of the discussion of normalisation. Further, it is argued that normalisers are better decomposed into activation-function-like maps with parameterised scaling. Overall, this constitutes a theoretically principled approach that yields new functions with empirical validation and raises questions about the affine + nonlinear approach.} }
Endnote
%0 Conference Paper %T The Affine Divergence: Aligning Activation Updates Beyond Normalisation %A George Bird %B Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling %C Proceedings of Machine Learning Research %D 2026 %E Alison Pouplin %E Sharvaree Vadgama %E Erik Bekkers %E Sékou-Oumar Kaba %E Hannah Lawrence %E Manuel Lecha %E Elizabeth Baker %E Julian Suk %E Robin Walters %E Jakub Tomczak %E Stefanie Jegelka %F pmlr-v326-bird26a %I PMLR %P 45--74 %U https://proceedings.mlr.press/v326/bird26a.html %V 326 %X A systematic mismatch exists between mathematically ideal and effective activation updates during gradient descent. As intended, parameters update in their direction of steepest descent. However, activations are argued to constitute a more directly impactful quantity to prioritise in optimisation, as they are closer to the loss in the computational graph and carry sample-dependent information through the network. Yet their propagated updates do not take the optimal steepest-descent step. These quantities exhibit non-ideal sample-wise scaling across affine, convolutional, and attention layers. Solutions to correct for this are trivial and, incidentally, derive normalisation from first principles despite motivational independence. Consequently, such considerations offer a fresh, conceptual reframe of normalisation’s action, with auxiliary experiments bolstering this mechanistic interpretation. Moreover, this analysis makes clear a second possibility: a solution that is functionally distinct from modern normalisations, without scale invariance, yet remains empirically successful — an alternative to the affine map. This outperforms conventional normalisers across several tests. This generalises to convolution via a new functional form, “PatchNorm”, a compositionally inseparable normaliser. Together, these provide an alternative mechanistic framework that both adds to and counters some of the discussion of normalisation. Further, it is argued that normalisers are better decomposed into activation-function-like maps with parameterised scaling. Overall, this constitutes a theoretically principled approach that yields new functions with empirical validation and raises questions about the affine + nonlinear approach.
APA
Bird, G.. (2026). The Affine Divergence: Aligning Activation Updates Beyond Normalisation. Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling, in Proceedings of Machine Learning Research 326:45-74 Available from https://proceedings.mlr.press/v326/bird26a.html.

Related Material