Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts Architectures

Venmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour, Tomer Asida, Abhinav Khattar, Nave Assaf, Maximilian Golub, Joseph Guman, Tiyasa Mitra, Ritchie Zhao, Ritika Borkar, Ran Zilberstein, Mostofa Patwary, Mohammad Shoeybi, Bita Darvish Rouhani
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27809-27821, 2026.

Abstract

Mixture-of-Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal for inference cost, as measured by accuracy per floating-point operation and per parameter. In this work, we revisit MoE design from a hardware-software co-design perspective, grounded in empirical and theoretical considerations. We characterize key performance bottlenecks across diverse deployment regimes, spanning offline high-throughput execution and online, latency-critical inference. Guided by these insights, we introduce LatentMoE, a new model architecture resulting from systematic design exploration and optimized for maximal accuracy per unit of compute. Empirical design space exploration at scales of up to 95B parameters and over a 1T-token training horizon, together with supporting theoretical analysis, shows that LatentMoE consistently outperforms standard MoE architectures in terms of accuracy per FLOP and per parameter. Given its strong performance, the LatentMoE architecture has been adopted by the flagship Nemotron-3 Super and Ultra models and scaled to substantially larger regimes, including longer token horizons and larger model sizes, as reported in (NVIDIA et al., 2025, arXiv:2512.20856).

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-elango26a, title = {Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts Architectures}, author = {Elango, Venmugil and Bhatia, Nidhi and Waleffe, Roger and Shafipour, Rasoul and Asida, Tomer and Khattar, Abhinav and Assaf, Nave and Golub, Maximilian and Guman, Joseph and Mitra, Tiyasa and Zhao, Ritchie and Borkar, Ritika and Zilberstein, Ran and Patwary, Mostofa and Shoeybi, Mohammad and Darvish Rouhani, Bita}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27809--27821}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/elango26a/elango26a.pdf}, url = {https://proceedings.mlr.press/v306/elango26a.html}, abstract = {Mixture-of-Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal for inference cost, as measured by accuracy per floating-point operation and per parameter. In this work, we revisit MoE design from a hardware-software co-design perspective, grounded in empirical and theoretical considerations. We characterize key performance bottlenecks across diverse deployment regimes, spanning offline high-throughput execution and online, latency-critical inference. Guided by these insights, we introduce LatentMoE, a new model architecture resulting from systematic design exploration and optimized for maximal accuracy per unit of compute. Empirical design space exploration at scales of up to 95B parameters and over a 1T-token training horizon, together with supporting theoretical analysis, shows that LatentMoE consistently outperforms standard MoE architectures in terms of accuracy per FLOP and per parameter. Given its strong performance, the LatentMoE architecture has been adopted by the flagship Nemotron-3 Super and Ultra models and scaled to substantially larger regimes, including longer token horizons and larger model sizes, as reported in (NVIDIA et al., 2025, arXiv:2512.20856).} }
Endnote
%0 Conference Paper %T Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts Architectures %A Venmugil Elango %A Nidhi Bhatia %A Roger Waleffe %A Rasoul Shafipour %A Tomer Asida %A Abhinav Khattar %A Nave Assaf %A Maximilian Golub %A Joseph Guman %A Tiyasa Mitra %A Ritchie Zhao %A Ritika Borkar %A Ran Zilberstein %A Mostofa Patwary %A Mohammad Shoeybi %A Bita Darvish Rouhani %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-elango26a %I PMLR %P 27809--27821 %U https://proceedings.mlr.press/v306/elango26a.html %V 306 %X Mixture-of-Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal for inference cost, as measured by accuracy per floating-point operation and per parameter. In this work, we revisit MoE design from a hardware-software co-design perspective, grounded in empirical and theoretical considerations. We characterize key performance bottlenecks across diverse deployment regimes, spanning offline high-throughput execution and online, latency-critical inference. Guided by these insights, we introduce LatentMoE, a new model architecture resulting from systematic design exploration and optimized for maximal accuracy per unit of compute. Empirical design space exploration at scales of up to 95B parameters and over a 1T-token training horizon, together with supporting theoretical analysis, shows that LatentMoE consistently outperforms standard MoE architectures in terms of accuracy per FLOP and per parameter. Given its strong performance, the LatentMoE architecture has been adopted by the flagship Nemotron-3 Super and Ultra models and scaled to substantially larger regimes, including longer token horizons and larger model sizes, as reported in (NVIDIA et al., 2025, arXiv:2512.20856).
APA
Elango, V., Bhatia, N., Waleffe, R., Shafipour, R., Asida, T., Khattar, A., Assaf, N., Golub, M., Guman, J., Mitra, T., Zhao, R., Borkar, R., Zilberstein, R., Patwary, M., Shoeybi, M. & Darvish Rouhani, B.. (2026). Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts Architectures. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27809-27821 Available from https://proceedings.mlr.press/v306/elango26a.html.

Related Material