ZipMoE: A Theoretically-Grounded Mixture of Experts Approach forParameter-Efficient Deep Learning

Lin Chen, Kyriakos Axiotis, Gang Fu, Kaiyuan Wang, Mohammadhossein Bateni, Vahab Mirrokni
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4096-4104, 2026.

Abstract

The relentless growth of large language models (LLMs) presents formidable challenges for their training and deployment. To address this critical bottleneck, we introduce ZipMoE, a novel family of parameter-efficient building blocks inspired by the Mixture of Experts (MoE) paradigm. ZipMoE offers a modular and highly efficient substitute for conventional fully connected layers. We provide a rigorous theoretical analysis of ZipMoE’s expressiveness, formally demonstrating its superior representational capacity over low-rank factorization. Furthermore, in a least squares regression setting, we prove that ZipMoE achieves a lower test error bound. Our empirical results—featuring comprehensive comparisons against low-rank, Monarch, and Kronecker methods—corroborate these theoretical findings. We demonstrate that ZipMoE consistently attains superior model quality under equivalent parameter or FLOP budgets, establishing it as a potent component for building efficient and powerful deep learning architectures.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-chen26f, title = { ZipMoE: A Theoretically-Grounded Mixture of Experts Approach forParameter-Efficient Deep Learning }, author = {Chen, Lin and Axiotis, Kyriakos and Fu, Gang and Wang, Kaiyuan and Bateni, Mohammadhossein and Mirrokni, Vahab}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4096--4104}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/chen26f/chen26f.pdf}, url = {https://proceedings.mlr.press/v300/chen26f.html}, abstract = { The relentless growth of large language models (LLMs) presents formidable challenges for their training and deployment. To address this critical bottleneck, we introduce ZipMoE, a novel family of parameter-efficient building blocks inspired by the Mixture of Experts (MoE) paradigm. ZipMoE offers a modular and highly efficient substitute for conventional fully connected layers. We provide a rigorous theoretical analysis of ZipMoE’s expressiveness, formally demonstrating its superior representational capacity over low-rank factorization. Furthermore, in a least squares regression setting, we prove that ZipMoE achieves a lower test error bound. Our empirical results—featuring comprehensive comparisons against low-rank, Monarch, and Kronecker methods—corroborate these theoretical findings. We demonstrate that ZipMoE consistently attains superior model quality under equivalent parameter or FLOP budgets, establishing it as a potent component for building efficient and powerful deep learning architectures. } }
Endnote
%0 Conference Paper %T ZipMoE: A Theoretically-Grounded Mixture of Experts Approach forParameter-Efficient Deep Learning %A Lin Chen %A Kyriakos Axiotis %A Gang Fu %A Kaiyuan Wang %A Mohammadhossein Bateni %A Vahab Mirrokni %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-chen26f %I PMLR %P 4096--4104 %U https://proceedings.mlr.press/v300/chen26f.html %V 300 %X The relentless growth of large language models (LLMs) presents formidable challenges for their training and deployment. To address this critical bottleneck, we introduce ZipMoE, a novel family of parameter-efficient building blocks inspired by the Mixture of Experts (MoE) paradigm. ZipMoE offers a modular and highly efficient substitute for conventional fully connected layers. We provide a rigorous theoretical analysis of ZipMoE’s expressiveness, formally demonstrating its superior representational capacity over low-rank factorization. Furthermore, in a least squares regression setting, we prove that ZipMoE achieves a lower test error bound. Our empirical results—featuring comprehensive comparisons against low-rank, Monarch, and Kronecker methods—corroborate these theoretical findings. We demonstrate that ZipMoE consistently attains superior model quality under equivalent parameter or FLOP budgets, establishing it as a potent component for building efficient and powerful deep learning architectures.
APA
Chen, L., Axiotis, K., Fu, G., Wang, K., Bateni, M. & Mirrokni, V.. (2026). ZipMoE: A Theoretically-Grounded Mixture of Experts Approach forParameter-Efficient Deep Learning . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4096-4104 Available from https://proceedings.mlr.press/v300/chen26f.html.

Related Material