Two mathematical models of knowledge distillation

Audrey Xie, Ludwig Schmidt, John Duchi
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4726-4734, 2026.

Abstract

Many hypotheses compete to explain the successes of knowledge distillation. To help address this, we propose and analyze a mathematical model of distillation, which suggests that distillation’s performance comes not from obtaining better models but from easier to optimize landscapes. For generalized linear models trained with stochastic gradient descent, we prove that distillation fits performant student models asymptotically more quickly than non-distilled models. In rank-1 matrix approximation, we characterize conditions on the target matrix under which gradient descent with distillation converges strictly faster than training on the supervised objective. The theory helps delineate the ways distillation provides benefits (i.e., in optimization speed, not in generalization), and experiments on real datasets corroborate the theoretical predictions.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-xie26b, title = { Two mathematical models of knowledge distillation }, author = {Xie, Audrey and Schmidt, Ludwig and Duchi, John}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4726--4734}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/xie26b/xie26b.pdf}, url = {https://proceedings.mlr.press/v300/xie26b.html}, abstract = { Many hypotheses compete to explain the successes of knowledge distillation. To help address this, we propose and analyze a mathematical model of distillation, which suggests that distillation’s performance comes not from obtaining better models but from easier to optimize landscapes. For generalized linear models trained with stochastic gradient descent, we prove that distillation fits performant student models asymptotically more quickly than non-distilled models. In rank-1 matrix approximation, we characterize conditions on the target matrix under which gradient descent with distillation converges strictly faster than training on the supervised objective. The theory helps delineate the ways distillation provides benefits (i.e., in optimization speed, not in generalization), and experiments on real datasets corroborate the theoretical predictions. } }
Endnote
%0 Conference Paper %T Two mathematical models of knowledge distillation %A Audrey Xie %A Ludwig Schmidt %A John Duchi %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-xie26b %I PMLR %P 4726--4734 %U https://proceedings.mlr.press/v300/xie26b.html %V 300 %X Many hypotheses compete to explain the successes of knowledge distillation. To help address this, we propose and analyze a mathematical model of distillation, which suggests that distillation’s performance comes not from obtaining better models but from easier to optimize landscapes. For generalized linear models trained with stochastic gradient descent, we prove that distillation fits performant student models asymptotically more quickly than non-distilled models. In rank-1 matrix approximation, we characterize conditions on the target matrix under which gradient descent with distillation converges strictly faster than training on the supervised objective. The theory helps delineate the ways distillation provides benefits (i.e., in optimization speed, not in generalization), and experiments on real datasets corroborate the theoretical predictions.
APA
Xie, A., Schmidt, L. & Duchi, J.. (2026). Two mathematical models of knowledge distillation . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4726-4734 Available from https://proceedings.mlr.press/v300/xie26b.html.

Related Material