$L_2$ Regularization for Learning Kernels

Corinna Cortes, Mehryar Mohri, Afshin Rostamizadeh
Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, PMLR R7:109-116, 2009.

Abstract

The choice of the kernel is critical to the success of many learning algorithms but it is typically left to the user. Instead, the training data can be used to learn the kernel by selecting it out of a given family, such as that of non-negativelinear combi- nations of p base kernels, constrained by a trace or L1 regularization. This paper studies the prob- lem of learning kernels with the same family of kernels but with an L2 regularization instead, and for regression problems. We analyze the prob- lem of learning kernels with ridge regression. We derive the form of the solution of the optimiza- tion problem and give an efficient iterative algo- rithm for computing that solution. We present a novel theoretical analysis of the problem based on stability and give learning bounds for orthog- onal kernels that contain only an additive term O( p p/m) when compared to the standard ker- nel ridge regression stability bound. We also re- port the results of experiments indicating that L1 regularization can lead to modest improvements for a small number of kernels, but to performance degradations in larger-scale cases.In contrast, L2 regularization never degrades performance and in fact achieves significant improvements with a large number of kernels.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR7-cortes09a, title = {$L_2$ Regularization for Learning Kernels}, author = {Cortes, Corinna and Mohri, Mehryar and Rostamizadeh, Afshin}, booktitle = {Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence}, pages = {109--116}, year = {2009}, editor = {Bilmes, Jeff and Ng, Andrew Y.}, volume = {R7}, series = {Proceedings of Machine Learning Research}, month = {18--21 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r7/main/assets/cortes09a/cortes09a.pdf}, url = {https://proceedings.mlr.press/r7/cortes09a.html}, abstract = {The choice of the kernel is critical to the success of many learning algorithms but it is typically left to the user. Instead, the training data can be used to learn the kernel by selecting it out of a given family, such as that of non-negativelinear combi- nations of p base kernels, constrained by a trace or L1 regularization. This paper studies the prob- lem of learning kernels with the same family of kernels but with an L2 regularization instead, and for regression problems. We analyze the prob- lem of learning kernels with ridge regression. We derive the form of the solution of the optimiza- tion problem and give an efficient iterative algo- rithm for computing that solution. We present a novel theoretical analysis of the problem based on stability and give learning bounds for orthog- onal kernels that contain only an additive term O( p p/m) when compared to the standard ker- nel ridge regression stability bound. We also re- port the results of experiments indicating that L1 regularization can lead to modest improvements for a small number of kernels, but to performance degradations in larger-scale cases.In contrast, L2 regularization never degrades performance and in fact achieves significant improvements with a large number of kernels.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T $L_2$ Regularization for Learning Kernels %A Corinna Cortes %A Mehryar Mohri %A Afshin Rostamizadeh %B Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2009 %E Jeff Bilmes %E Andrew Y. Ng %F pmlr-vR7-cortes09a %I PMLR %P 109--116 %U https://proceedings.mlr.press/r7/cortes09a.html %V R7 %X The choice of the kernel is critical to the success of many learning algorithms but it is typically left to the user. Instead, the training data can be used to learn the kernel by selecting it out of a given family, such as that of non-negativelinear combi- nations of p base kernels, constrained by a trace or L1 regularization. This paper studies the prob- lem of learning kernels with the same family of kernels but with an L2 regularization instead, and for regression problems. We analyze the prob- lem of learning kernels with ridge regression. We derive the form of the solution of the optimiza- tion problem and give an efficient iterative algo- rithm for computing that solution. We present a novel theoretical analysis of the problem based on stability and give learning bounds for orthog- onal kernels that contain only an additive term O( p p/m) when compared to the standard ker- nel ridge regression stability bound. We also re- port the results of experiments indicating that L1 regularization can lead to modest improvements for a small number of kernels, but to performance degradations in larger-scale cases.In contrast, L2 regularization never degrades performance and in fact achieves significant improvements with a large number of kernels. %Z Reissued by PMLR on 04 October 2026.
APA
Cortes, C., Mohri, M. & Rostamizadeh, A.. (2009). $L_2$ Regularization for Learning Kernels. Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R7:109-116 Available from https://proceedings.mlr.press/r7/cortes09a.html. Reissued by PMLR on 04 October 2026.

Related Material