DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts

Jiarui Feng, Hanqing Zeng, Karish Grover, Ruizhong Qiu, Yinglong Xia, Qiang Zhang, Qifan Wang, Ren Chen, Dongqi Fu, Jiayi Liu, Zhuokai Zhao, Xiangjun Fan, Benyu Zhang, Yixin Chen
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:30510-30530, 2026.

Abstract

Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge. Prior work shows that fine-grained experts enlarge the space of expert combinations and improve flexibility, but they also impose substantial routing overhead, creating a new scalability bottleneck. In this paper, we explore a complementary axis for scaling—how expert outputs are aggregated. We theoretically show that replacing the standard weighted-summation aggregation with structural aggregation expands the expert-combination space without altering the experts or router, and enables possible multi-step reasoning within a single MoE layer. To this end, we propose DAG-MoE, a sparse MoE framework that employs a lightweight module to automatically learn the optimal aggregation structure among the selected experts. Extensive experiments under standard language modeling settings show that DAG-MoE consistently improves performance in both pretraining and fine-tuning, surpassing traditional MoE baselines.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-feng26v, title = {{DAG}-{M}o{E}: From Simple Mixture to Structural Aggregation in Mixture-of-Experts}, author = {Feng, Jiarui and Zeng, Hanqing and Grover, Karish and Qiu, Ruizhong and Xia, Yinglong and Zhang, Qiang and Wang, Qifan and Chen, Ren and Fu, Dongqi and Liu, Jiayi and Zhao, Zhuokai and Fan, Xiangjun and Zhang, Benyu and Chen, Yixin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {30510--30530}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/feng26v/feng26v.pdf}, url = {https://proceedings.mlr.press/v306/feng26v.html}, abstract = {Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge. Prior work shows that fine-grained experts enlarge the space of expert combinations and improve flexibility, but they also impose substantial routing overhead, creating a new scalability bottleneck. In this paper, we explore a complementary axis for scaling—how expert outputs are aggregated. We theoretically show that replacing the standard weighted-summation aggregation with structural aggregation expands the expert-combination space without altering the experts or router, and enables possible multi-step reasoning within a single MoE layer. To this end, we propose DAG-MoE, a sparse MoE framework that employs a lightweight module to automatically learn the optimal aggregation structure among the selected experts. Extensive experiments under standard language modeling settings show that DAG-MoE consistently improves performance in both pretraining and fine-tuning, surpassing traditional MoE baselines.} }
Endnote
%0 Conference Paper %T DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts %A Jiarui Feng %A Hanqing Zeng %A Karish Grover %A Ruizhong Qiu %A Yinglong Xia %A Qiang Zhang %A Qifan Wang %A Ren Chen %A Dongqi Fu %A Jiayi Liu %A Zhuokai Zhao %A Xiangjun Fan %A Benyu Zhang %A Yixin Chen %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-feng26v %I PMLR %P 30510--30530 %U https://proceedings.mlr.press/v306/feng26v.html %V 306 %X Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge. Prior work shows that fine-grained experts enlarge the space of expert combinations and improve flexibility, but they also impose substantial routing overhead, creating a new scalability bottleneck. In this paper, we explore a complementary axis for scaling—how expert outputs are aggregated. We theoretically show that replacing the standard weighted-summation aggregation with structural aggregation expands the expert-combination space without altering the experts or router, and enables possible multi-step reasoning within a single MoE layer. To this end, we propose DAG-MoE, a sparse MoE framework that employs a lightweight module to automatically learn the optimal aggregation structure among the selected experts. Extensive experiments under standard language modeling settings show that DAG-MoE consistently improves performance in both pretraining and fine-tuning, surpassing traditional MoE baselines.
APA
Feng, J., Zeng, H., Grover, K., Qiu, R., Xia, Y., Zhang, Q., Wang, Q., Chen, R., Fu, D., Liu, J., Zhao, Z., Fan, X., Zhang, B. & Chen, Y.. (2026). DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:30510-30530 Available from https://proceedings.mlr.press/v306/feng26v.html.

Related Material