Diversity-Aware Recursive Feature Multiple Kernel Learning

Nan Cao, Xu Zhao, Teng Zhang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:11815-11829, 2026.

Abstract

Multiple kernel learning (MKL) combines several base kernels in the spirit of ensemble learning, yet existing methods rarely model kernel diversity—a known cornerstone of ensembles—and most traditional kernels weight all features uniformly, ignoring feature-level discriminability. We address both gaps with DARFMMKL: a data-driven kernel family (Recursive Feature Machine kernels) that learns feature importance directly from data, paired with a kernel selection method that jointly optimizes diversity and quality. The resulting NP-hard binary quadratic program is reformulated via Glover linearization and continuous relaxation into a linear program, and accelerated by Nyström sketching, yielding a selector whose cost is decoupled from the sample size. We provide a covering-number generalization bound that explicitly relates kernel diversity to estimation error. Experiments on 12 benchmark datasets show that DARFMMKL consistently outperforms 9 state-of-the-art MKL methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cao26ae, title = {Diversity-Aware Recursive Feature Multiple Kernel Learning}, author = {Cao, Nan and Zhao, Xu and Zhang, Teng}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {11815--11829}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cao26ae/cao26ae.pdf}, url = {https://proceedings.mlr.press/v306/cao26ae.html}, abstract = {Multiple kernel learning (MKL) combines several base kernels in the spirit of ensemble learning, yet existing methods rarely model kernel diversity—a known cornerstone of ensembles—and most traditional kernels weight all features uniformly, ignoring feature-level discriminability. We address both gaps with DARFMMKL: a data-driven kernel family (Recursive Feature Machine kernels) that learns feature importance directly from data, paired with a kernel selection method that jointly optimizes diversity and quality. The resulting NP-hard binary quadratic program is reformulated via Glover linearization and continuous relaxation into a linear program, and accelerated by Nyström sketching, yielding a selector whose cost is decoupled from the sample size. We provide a covering-number generalization bound that explicitly relates kernel diversity to estimation error. Experiments on 12 benchmark datasets show that DARFMMKL consistently outperforms 9 state-of-the-art MKL methods.} }
Endnote
%0 Conference Paper %T Diversity-Aware Recursive Feature Multiple Kernel Learning %A Nan Cao %A Xu Zhao %A Teng Zhang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cao26ae %I PMLR %P 11815--11829 %U https://proceedings.mlr.press/v306/cao26ae.html %V 306 %X Multiple kernel learning (MKL) combines several base kernels in the spirit of ensemble learning, yet existing methods rarely model kernel diversity—a known cornerstone of ensembles—and most traditional kernels weight all features uniformly, ignoring feature-level discriminability. We address both gaps with DARFMMKL: a data-driven kernel family (Recursive Feature Machine kernels) that learns feature importance directly from data, paired with a kernel selection method that jointly optimizes diversity and quality. The resulting NP-hard binary quadratic program is reformulated via Glover linearization and continuous relaxation into a linear program, and accelerated by Nyström sketching, yielding a selector whose cost is decoupled from the sample size. We provide a covering-number generalization bound that explicitly relates kernel diversity to estimation error. Experiments on 12 benchmark datasets show that DARFMMKL consistently outperforms 9 state-of-the-art MKL methods.
APA
Cao, N., Zhao, X. & Zhang, T.. (2026). Diversity-Aware Recursive Feature Multiple Kernel Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:11815-11829 Available from https://proceedings.mlr.press/v306/cao26ae.html.

Related Material