SpanNorm: Reconciling Training Stability and Performance in Deep Transformers

Wang Chao, Bei Li, Jiaqi Zhang, Xinyu Liu, Yuchun Fan, Linkun Lyu, Xin Chen, Jingang Wang, Tong Xiao, Peng Pei, Xunliang Cai
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:12948-12966, 2026.

Abstract

The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ”PreNorm” architecture ensures training stability at the cost of potential performance degradation in deep models, while the ”PostNorm” architecture offers strong performance but suffers from severe training instability. In this work, we propose SpanNorm, a novel technique designed to resolve this dilemma by integrating the strengths of both paradigms. SpanNorm adopts the clean residual path of PreNorm to stabilize signal propagation while employing a PostNorm-style computation that normalizes the output of the residual connection, thereby enhancing model performance. We provide a theoretical analysis demonstrating that SpanNorm, combined with a principled scaling strategy, maintains bounded signal variance throughout the network, preventing the gradient issues that plague PostNorm models, and alleviating the representation collapse of PreNorm. Empirically, SpanNorm consistently outperforms standard normalization schemes in both dense and Mixture-of-Experts (MoE) scenarios, paving the way for more powerful and stable Transformer architectures.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chao26b, title = {{S}pan{N}orm: Reconciling Training Stability and Performance in Deep Transformers}, author = {Chao, Wang and Li, Bei and Zhang, Jiaqi and Liu, Xinyu and Fan, Yuchun and Lyu, Linkun and Chen, Xin and Wang, Jingang and Xiao, Tong and Pei, Peng and Cai, Xunliang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {12948--12966}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chao26b/chao26b.pdf}, url = {https://proceedings.mlr.press/v306/chao26b.html}, abstract = {The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ”PreNorm” architecture ensures training stability at the cost of potential performance degradation in deep models, while the ”PostNorm” architecture offers strong performance but suffers from severe training instability. In this work, we propose SpanNorm, a novel technique designed to resolve this dilemma by integrating the strengths of both paradigms. SpanNorm adopts the clean residual path of PreNorm to stabilize signal propagation while employing a PostNorm-style computation that normalizes the output of the residual connection, thereby enhancing model performance. We provide a theoretical analysis demonstrating that SpanNorm, combined with a principled scaling strategy, maintains bounded signal variance throughout the network, preventing the gradient issues that plague PostNorm models, and alleviating the representation collapse of PreNorm. Empirically, SpanNorm consistently outperforms standard normalization schemes in both dense and Mixture-of-Experts (MoE) scenarios, paving the way for more powerful and stable Transformer architectures.} }
Endnote
%0 Conference Paper %T SpanNorm: Reconciling Training Stability and Performance in Deep Transformers %A Wang Chao %A Bei Li %A Jiaqi Zhang %A Xinyu Liu %A Yuchun Fan %A Linkun Lyu %A Xin Chen %A Jingang Wang %A Tong Xiao %A Peng Pei %A Xunliang Cai %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chao26b %I PMLR %P 12948--12966 %U https://proceedings.mlr.press/v306/chao26b.html %V 306 %X The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ”PreNorm” architecture ensures training stability at the cost of potential performance degradation in deep models, while the ”PostNorm” architecture offers strong performance but suffers from severe training instability. In this work, we propose SpanNorm, a novel technique designed to resolve this dilemma by integrating the strengths of both paradigms. SpanNorm adopts the clean residual path of PreNorm to stabilize signal propagation while employing a PostNorm-style computation that normalizes the output of the residual connection, thereby enhancing model performance. We provide a theoretical analysis demonstrating that SpanNorm, combined with a principled scaling strategy, maintains bounded signal variance throughout the network, preventing the gradient issues that plague PostNorm models, and alleviating the representation collapse of PreNorm. Empirically, SpanNorm consistently outperforms standard normalization schemes in both dense and Mixture-of-Experts (MoE) scenarios, paving the way for more powerful and stable Transformer architectures.
APA
Chao, W., Li, B., Zhang, J., Liu, X., Fan, Y., Lyu, L., Chen, X., Wang, J., Xiao, T., Pei, P. & Cai, X.. (2026). SpanNorm: Reconciling Training Stability and Performance in Deep Transformers. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:12948-12966 Available from https://proceedings.mlr.press/v306/chao26b.html.

Related Material