Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts

Fangshuo Liao, Anastasios Kyrillidis
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3898-3906, 2026.

Abstract

Mixture-of-Experts (MoE) architectures have emerged as a cornerstone of modern AI systems. In particular, MoEs route inputs dynamically to specialized experts, whose outputs are aggregated through weighted summation. Despite their widespread application, theoretical understanding of MoE training dynamics remains limited to either separate expert-router optimization or restrictive top-1 routing scenarios with carefully constructed datasets. This paper advances MoE theory by providing convergence guarantees for joint training of soft-routed MoE models with non-linear routers and experts in a student-teacher framework. We prove that, with moderate over-parameterization, the student network undergoes a feature learning phase, where the router’s learning process are “guided" by the experts, that recovers the teacher’s parameters. Moreover, we show that a post-training pruning can effectively eliminate redundant neurons, followed by a provably convergent fine-tuning process that reaches global optimality. Our analysis brings novel insight in understanding the optimization landscape of the MoE architecture.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-liao26a, title = { Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts }, author = {Liao, Fangshuo and Kyrillidis, Anastasios}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3898--3906}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/liao26a/liao26a.pdf}, url = {https://proceedings.mlr.press/v300/liao26a.html}, abstract = { Mixture-of-Experts (MoE) architectures have emerged as a cornerstone of modern AI systems. In particular, MoEs route inputs dynamically to specialized experts, whose outputs are aggregated through weighted summation. Despite their widespread application, theoretical understanding of MoE training dynamics remains limited to either separate expert-router optimization or restrictive top-1 routing scenarios with carefully constructed datasets. This paper advances MoE theory by providing convergence guarantees for joint training of soft-routed MoE models with non-linear routers and experts in a student-teacher framework. We prove that, with moderate over-parameterization, the student network undergoes a feature learning phase, where the router’s learning process are “guided" by the experts, that recovers the teacher’s parameters. Moreover, we show that a post-training pruning can effectively eliminate redundant neurons, followed by a provably convergent fine-tuning process that reaches global optimality. Our analysis brings novel insight in understanding the optimization landscape of the MoE architecture. } }
Endnote
%0 Conference Paper %T Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts %A Fangshuo Liao %A Anastasios Kyrillidis %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-liao26a %I PMLR %P 3898--3906 %U https://proceedings.mlr.press/v300/liao26a.html %V 300 %X Mixture-of-Experts (MoE) architectures have emerged as a cornerstone of modern AI systems. In particular, MoEs route inputs dynamically to specialized experts, whose outputs are aggregated through weighted summation. Despite their widespread application, theoretical understanding of MoE training dynamics remains limited to either separate expert-router optimization or restrictive top-1 routing scenarios with carefully constructed datasets. This paper advances MoE theory by providing convergence guarantees for joint training of soft-routed MoE models with non-linear routers and experts in a student-teacher framework. We prove that, with moderate over-parameterization, the student network undergoes a feature learning phase, where the router’s learning process are “guided" by the experts, that recovers the teacher’s parameters. Moreover, we show that a post-training pruning can effectively eliminate redundant neurons, followed by a provably convergent fine-tuning process that reaches global optimality. Our analysis brings novel insight in understanding the optimization landscape of the MoE architecture.
APA
Liao, F. & Kyrillidis, A.. (2026). Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3898-3906 Available from https://proceedings.mlr.press/v300/liao26a.html.

Related Material