SimpleGPT: Improving GPT via A Simple Normalization Strategy

Marco Chen, Xianbiao Qi, Yelin He, Jiaquan Ye, Rong Xiao
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16361-16386, 2026.

Abstract

In this work, we revisit Transformer optimization through the lens of second-order geometry and establish a direct connection between architectural design, activation scale, the Hessian matrix, and the maximum tolerable learning rate. We introduce a simple normalization strategy, termed SimpleNorm, which stabilizes intermediate activation scales by construction. Then, by analyzing the Hessian of the loss with respect to network activations, we theoretically show that SimpleNorm significantly reduces the spectral norm of the Hessian, thereby permitting larger stable learning rates. We validate our theoretical findings through extensive experiments on large GPT models at parameter scales 1B, 1.4B, 7B and 8B. Empirically, SimpleGPT, our SimpleNorm-based network, tolerates learning rates 3$\times$-10$\times$ larger than standard convention, consistently demonstrates strong optimization stability, and achieves substantially better performance than well-established baselines. Specifically, when training 7B-scale models for 60K steps, SimpleGPT reduces the training loss from 2.290 to 2.208 compared with Llama2 with QKNorm. Our code is available at https://github.com/Ocram7/SimpleGPT.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26dj, title = {{S}imple{GPT}: Improving {GPT} via A Simple Normalization Strategy}, author = {Chen, Marco and Qi, Xianbiao and He, Yelin and Ye, Jiaquan and Xiao, Rong}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16361--16386}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26dj/chen26dj.pdf}, url = {https://proceedings.mlr.press/v306/chen26dj.html}, abstract = {In this work, we revisit Transformer optimization through the lens of second-order geometry and establish a direct connection between architectural design, activation scale, the Hessian matrix, and the maximum tolerable learning rate. We introduce a simple normalization strategy, termed SimpleNorm, which stabilizes intermediate activation scales by construction. Then, by analyzing the Hessian of the loss with respect to network activations, we theoretically show that SimpleNorm significantly reduces the spectral norm of the Hessian, thereby permitting larger stable learning rates. We validate our theoretical findings through extensive experiments on large GPT models at parameter scales 1B, 1.4B, 7B and 8B. Empirically, SimpleGPT, our SimpleNorm-based network, tolerates learning rates 3$\times$-10$\times$ larger than standard convention, consistently demonstrates strong optimization stability, and achieves substantially better performance than well-established baselines. Specifically, when training 7B-scale models for 60K steps, SimpleGPT reduces the training loss from 2.290 to 2.208 compared with Llama2 with QKNorm. Our code is available at https://github.com/Ocram7/SimpleGPT.} }
Endnote
%0 Conference Paper %T SimpleGPT: Improving GPT via A Simple Normalization Strategy %A Marco Chen %A Xianbiao Qi %A Yelin He %A Jiaquan Ye %A Rong Xiao %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26dj %I PMLR %P 16361--16386 %U https://proceedings.mlr.press/v306/chen26dj.html %V 306 %X In this work, we revisit Transformer optimization through the lens of second-order geometry and establish a direct connection between architectural design, activation scale, the Hessian matrix, and the maximum tolerable learning rate. We introduce a simple normalization strategy, termed SimpleNorm, which stabilizes intermediate activation scales by construction. Then, by analyzing the Hessian of the loss with respect to network activations, we theoretically show that SimpleNorm significantly reduces the spectral norm of the Hessian, thereby permitting larger stable learning rates. We validate our theoretical findings through extensive experiments on large GPT models at parameter scales 1B, 1.4B, 7B and 8B. Empirically, SimpleGPT, our SimpleNorm-based network, tolerates learning rates 3$\times$-10$\times$ larger than standard convention, consistently demonstrates strong optimization stability, and achieves substantially better performance than well-established baselines. Specifically, when training 7B-scale models for 60K steps, SimpleGPT reduces the training loss from 2.290 to 2.208 compared with Llama2 with QKNorm. Our code is available at https://github.com/Ocram7/SimpleGPT.
APA
Chen, M., Qi, X., He, Y., Ye, J. & Xiao, R.. (2026). SimpleGPT: Improving GPT via A Simple Normalization Strategy. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16361-16386 Available from https://proceedings.mlr.press/v306/chen26dj.html.

Related Material