On the Weight Density of L2-Regularized Linear Classification and Regression

He-Zhe Lin, Zhi-Bao Lu, Sheng-Wei Chen, Cheng-Hung Liu, Chih-Jen Lin
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:127-135, 2026.

Abstract

For traditional linear models with the widely used $L_2$-regularizer, it is often assumed that the resulting models are dense. As a result, little attention has been paid to when the optimal solution for an $L_2$-regularized problem can actually be sparse. In this work, we rigorously prove that for $L_2$-regularized support vector classification/regression, the theoretical optimum can indeed be sparse when the data have sparse feature values. Surprisingly, we observe that some optimization methods fail to preserve this sparsity and instead produce fully dense numerical solutions, leading to unnecessary storage overhead. We explain this phenomenon through detailed analysis. In particular, we novelly show that certain coordinate descent methods naturally yields sparser numerical solutions compared to other optimization algorithms. By applying suitable algorithms that preserve numerical sparsity, the storage can be reduced by up to 50%, which is highly advantageous for large-scale industrial applications.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-lin26a, title = { On the Weight Density of L2-Regularized Linear Classification and Regression }, author = {Lin, He-Zhe and Lu, Zhi-Bao and Chen, Sheng-Wei and Liu, Cheng-Hung and Lin, Chih-Jen}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {127--135}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/lin26a/lin26a.pdf}, url = {https://proceedings.mlr.press/v300/lin26a.html}, abstract = { For traditional linear models with the widely used $L_2$-regularizer, it is often assumed that the resulting models are dense. As a result, little attention has been paid to when the optimal solution for an $L_2$-regularized problem can actually be sparse. In this work, we rigorously prove that for $L_2$-regularized support vector classification/regression, the theoretical optimum can indeed be sparse when the data have sparse feature values. Surprisingly, we observe that some optimization methods fail to preserve this sparsity and instead produce fully dense numerical solutions, leading to unnecessary storage overhead. We explain this phenomenon through detailed analysis. In particular, we novelly show that certain coordinate descent methods naturally yields sparser numerical solutions compared to other optimization algorithms. By applying suitable algorithms that preserve numerical sparsity, the storage can be reduced by up to 50%, which is highly advantageous for large-scale industrial applications. } }
Endnote
%0 Conference Paper %T On the Weight Density of L2-Regularized Linear Classification and Regression %A He-Zhe Lin %A Zhi-Bao Lu %A Sheng-Wei Chen %A Cheng-Hung Liu %A Chih-Jen Lin %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-lin26a %I PMLR %P 127--135 %U https://proceedings.mlr.press/v300/lin26a.html %V 300 %X For traditional linear models with the widely used $L_2$-regularizer, it is often assumed that the resulting models are dense. As a result, little attention has been paid to when the optimal solution for an $L_2$-regularized problem can actually be sparse. In this work, we rigorously prove that for $L_2$-regularized support vector classification/regression, the theoretical optimum can indeed be sparse when the data have sparse feature values. Surprisingly, we observe that some optimization methods fail to preserve this sparsity and instead produce fully dense numerical solutions, leading to unnecessary storage overhead. We explain this phenomenon through detailed analysis. In particular, we novelly show that certain coordinate descent methods naturally yields sparser numerical solutions compared to other optimization algorithms. By applying suitable algorithms that preserve numerical sparsity, the storage can be reduced by up to 50%, which is highly advantageous for large-scale industrial applications.
APA
Lin, H., Lu, Z., Chen, S., Liu, C. & Lin, C.. (2026). On the Weight Density of L2-Regularized Linear Classification and Regression . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:127-135 Available from https://proceedings.mlr.press/v300/lin26a.html.

Related Material