SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training

Mohammed Adnan, Rohan Jain, Tom Jacobs, Ekansh Sharma, Rahul G Krishnan, Rebekka Burkholz, Yani Ioannou
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:645-674, 2026.

Abstract

Dynamic Sparse Training (DST) methods train neural networks by maintaining sparsity while dynamically adapting the network topology. Despite the promise of reduced computation, DST methods converge significantly slower than dense training, often requiring comparable training time to achieve similar accuracy. We demonstrate both analytically and empirically that Batch Normalization (BN) adversely affects sparse training, and propose SparseOpt — a sparsity-aware optimizer — to address this. Experiments on ResNet models across CIFAR-100 and ImageNet demonstrate consistently faster convergence and improved generalization with our proposed method. Our work highlights the limitations of current normalization layers in sparse training and provides the first systematic study of the interaction between Batch Normalization, sparse layers, and DST, taking a significant step toward making DST practically competitive with dense training.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-adnan26a, title = {{S}parse{O}pt: Addressing Normalization-induced Gradient Skew in Sparse Training}, author = {Adnan, Mohammed and Jain, Rohan and Jacobs, Tom and Sharma, Ekansh and Krishnan, Rahul G and Burkholz, Rebekka and Ioannou, Yani}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {645--674}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/adnan26a/adnan26a.pdf}, url = {https://proceedings.mlr.press/v306/adnan26a.html}, abstract = {Dynamic Sparse Training (DST) methods train neural networks by maintaining sparsity while dynamically adapting the network topology. Despite the promise of reduced computation, DST methods converge significantly slower than dense training, often requiring comparable training time to achieve similar accuracy. We demonstrate both analytically and empirically that Batch Normalization (BN) adversely affects sparse training, and propose SparseOpt — a sparsity-aware optimizer — to address this. Experiments on ResNet models across CIFAR-100 and ImageNet demonstrate consistently faster convergence and improved generalization with our proposed method. Our work highlights the limitations of current normalization layers in sparse training and provides the first systematic study of the interaction between Batch Normalization, sparse layers, and DST, taking a significant step toward making DST practically competitive with dense training.} }
Endnote
%0 Conference Paper %T SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training %A Mohammed Adnan %A Rohan Jain %A Tom Jacobs %A Ekansh Sharma %A Rahul G Krishnan %A Rebekka Burkholz %A Yani Ioannou %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-adnan26a %I PMLR %P 645--674 %U https://proceedings.mlr.press/v306/adnan26a.html %V 306 %X Dynamic Sparse Training (DST) methods train neural networks by maintaining sparsity while dynamically adapting the network topology. Despite the promise of reduced computation, DST methods converge significantly slower than dense training, often requiring comparable training time to achieve similar accuracy. We demonstrate both analytically and empirically that Batch Normalization (BN) adversely affects sparse training, and propose SparseOpt — a sparsity-aware optimizer — to address this. Experiments on ResNet models across CIFAR-100 and ImageNet demonstrate consistently faster convergence and improved generalization with our proposed method. Our work highlights the limitations of current normalization layers in sparse training and provides the first systematic study of the interaction between Batch Normalization, sparse layers, and DST, taking a significant step toward making DST practically competitive with dense training.
APA
Adnan, M., Jain, R., Jacobs, T., Sharma, E., Krishnan, R.G., Burkholz, R. & Ioannou, Y.. (2026). SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:645-674 Available from https://proceedings.mlr.press/v306/adnan26a.html.

Related Material