Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism

Chenwei Cui, Rockwell Jackson, Benjamin Joseph Herrera, Ana M. Tárano, Hannah Kerner
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:21847-21863, 2026.

Abstract

Large language models have transformed many applications but remain expensive to train. Sparse Mixture of Experts (MoE) addresses this through conditional computation, with Expert Parallel (EP) as the standard distributed training method. However, EP has three limitations: communication cost grows linearly with the number of activated experts $k$, load imbalance affects latency and memory usage, and data-dependent communication requires metadata exchange. We propose Multi-Head LatentMoE and Head Parallel (HP), a new architecture and parallelism that achieve $O(1)$ communication cost regardless of $k$, completely balanced traffic, and deterministic communication, all while remaining compatible with EP. To accelerate Multi-Head LatentMoE, we propose IO-aware routing and expert computation. Compared to MoE with EP, Multi-Head LatentMoE with HP trains up to $1.82\times$ faster while having better performance. With double the granularity, the performance is even better while being $1.08\times$ faster. Our method makes multi-billion-parameter foundation model research more accessible.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cui26c, title = {Multi-Head {L}atent{M}o{E} and Head Parallel: Communication-Efficient and Deterministic {M}o{E} Parallelism}, author = {Cui, Chenwei and Jackson, Rockwell and Herrera, Benjamin Joseph and T\'{a}rano, Ana M. and Kerner, Hannah}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {21847--21863}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cui26c/cui26c.pdf}, url = {https://proceedings.mlr.press/v306/cui26c.html}, abstract = {Large language models have transformed many applications but remain expensive to train. Sparse Mixture of Experts (MoE) addresses this through conditional computation, with Expert Parallel (EP) as the standard distributed training method. However, EP has three limitations: communication cost grows linearly with the number of activated experts $k$, load imbalance affects latency and memory usage, and data-dependent communication requires metadata exchange. We propose Multi-Head LatentMoE and Head Parallel (HP), a new architecture and parallelism that achieve $O(1)$ communication cost regardless of $k$, completely balanced traffic, and deterministic communication, all while remaining compatible with EP. To accelerate Multi-Head LatentMoE, we propose IO-aware routing and expert computation. Compared to MoE with EP, Multi-Head LatentMoE with HP trains up to $1.82\times$ faster while having better performance. With double the granularity, the performance is even better while being $1.08\times$ faster. Our method makes multi-billion-parameter foundation model research more accessible.} }
Endnote
%0 Conference Paper %T Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism %A Chenwei Cui %A Rockwell Jackson %A Benjamin Joseph Herrera %A Ana M. Tárano %A Hannah Kerner %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cui26c %I PMLR %P 21847--21863 %U https://proceedings.mlr.press/v306/cui26c.html %V 306 %X Large language models have transformed many applications but remain expensive to train. Sparse Mixture of Experts (MoE) addresses this through conditional computation, with Expert Parallel (EP) as the standard distributed training method. However, EP has three limitations: communication cost grows linearly with the number of activated experts $k$, load imbalance affects latency and memory usage, and data-dependent communication requires metadata exchange. We propose Multi-Head LatentMoE and Head Parallel (HP), a new architecture and parallelism that achieve $O(1)$ communication cost regardless of $k$, completely balanced traffic, and deterministic communication, all while remaining compatible with EP. To accelerate Multi-Head LatentMoE, we propose IO-aware routing and expert computation. Compared to MoE with EP, Multi-Head LatentMoE with HP trains up to $1.82\times$ faster while having better performance. With double the granularity, the performance is even better while being $1.08\times$ faster. Our method makes multi-billion-parameter foundation model research more accessible.
APA
Cui, C., Jackson, R., Herrera, B.J., Tárano, A.M. & Kerner, H.. (2026). Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:21847-21863 Available from https://proceedings.mlr.press/v306/cui26c.html.

Related Material