Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers

Anrui Chen, Ruijun Huang, Xin Zhang, Fang Dong, Hengjie Cao, Zhendong Huang, Yifeng Yang, Mengyi Chen, Jixian Zhou, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Tun Lu, Fan Yang, Li Shang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:14688-14706, 2026.

Abstract

Mixture-of-Experts (MoE) architectures are appealing for continual learning because sparse routing should localize updates and reduce interference, yet MoE Transformers still forget substantially even with sparse, well-balanced expert utilization. We attribute this gap to a pre-routing bottleneck: multi-head attention concatenates head-specific signals into a single post-attention router input, forcing routing to act on co-occurring feature compositions rather than separable head channels. We show that this router input simultaneously encodes multiple separately decodable semantic and structural factors with uneven head support, and that different feature compositions induce weakly aligned parameter-gradient directions; as a result, routing maps many distinct compositions to the same route. We quantify this collision effect via a route-wise effective composition number $N_{\mathrm{eff}}$ and find that higher $N_{\mathrm{eff}}$ is associated with larger old-task loss increases after continual training. Motivated by these findings, we propose MH-MoE, which performs head-wise routing over sub-representations to increase routing granularity and reduce composition collisions. On TRACE across multiple backbones, MH-MoE consistently improves the retention–accuracy trade-off over LoRA-MoE variants.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26aw, title = {Multi-Head Attention as a Source of Catastrophic Forgetting in {M}o{E} Transformers}, author = {Chen, Anrui and Huang, Ruijun and Zhang, Xin and Dong, Fang and Cao, Hengjie and Huang, Zhendong and Yang, Yifeng and Chen, Mengyi and Zhou, Jixian and Dong, Mingzhi and Wang, Yujiang and Hou, Jinlong and Lv, Qin and Dick, Robert P. and Cheng, Yuan and Lu, Tun and Yang, Fan and Shang, Li}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {14688--14706}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26aw/chen26aw.pdf}, url = {https://proceedings.mlr.press/v306/chen26aw.html}, abstract = {Mixture-of-Experts (MoE) architectures are appealing for continual learning because sparse routing should localize updates and reduce interference, yet MoE Transformers still forget substantially even with sparse, well-balanced expert utilization. We attribute this gap to a pre-routing bottleneck: multi-head attention concatenates head-specific signals into a single post-attention router input, forcing routing to act on co-occurring feature compositions rather than separable head channels. We show that this router input simultaneously encodes multiple separately decodable semantic and structural factors with uneven head support, and that different feature compositions induce weakly aligned parameter-gradient directions; as a result, routing maps many distinct compositions to the same route. We quantify this collision effect via a route-wise effective composition number $N_{\mathrm{eff}}$ and find that higher $N_{\mathrm{eff}}$ is associated with larger old-task loss increases after continual training. Motivated by these findings, we propose MH-MoE, which performs head-wise routing over sub-representations to increase routing granularity and reduce composition collisions. On TRACE across multiple backbones, MH-MoE consistently improves the retention–accuracy trade-off over LoRA-MoE variants.} }
Endnote
%0 Conference Paper %T Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers %A Anrui Chen %A Ruijun Huang %A Xin Zhang %A Fang Dong %A Hengjie Cao %A Zhendong Huang %A Yifeng Yang %A Mengyi Chen %A Jixian Zhou %A Mingzhi Dong %A Yujiang Wang %A Jinlong Hou %A Qin Lv %A Robert P. Dick %A Yuan Cheng %A Tun Lu %A Fan Yang %A Li Shang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26aw %I PMLR %P 14688--14706 %U https://proceedings.mlr.press/v306/chen26aw.html %V 306 %X Mixture-of-Experts (MoE) architectures are appealing for continual learning because sparse routing should localize updates and reduce interference, yet MoE Transformers still forget substantially even with sparse, well-balanced expert utilization. We attribute this gap to a pre-routing bottleneck: multi-head attention concatenates head-specific signals into a single post-attention router input, forcing routing to act on co-occurring feature compositions rather than separable head channels. We show that this router input simultaneously encodes multiple separately decodable semantic and structural factors with uneven head support, and that different feature compositions induce weakly aligned parameter-gradient directions; as a result, routing maps many distinct compositions to the same route. We quantify this collision effect via a route-wise effective composition number $N_{\mathrm{eff}}$ and find that higher $N_{\mathrm{eff}}$ is associated with larger old-task loss increases after continual training. Motivated by these findings, we propose MH-MoE, which performs head-wise routing over sub-representations to increase routing granularity and reduce composition collisions. On TRACE across multiple backbones, MH-MoE consistently improves the retention–accuracy trade-off over LoRA-MoE variants.
APA
Chen, A., Huang, R., Zhang, X., Dong, F., Cao, H., Huang, Z., Yang, Y., Chen, M., Zhou, J., Dong, M., Wang, Y., Hou, J., Lv, Q., Dick, R.P., Cheng, Y., Lu, T., Yang, F. & Shang, L.. (2026). Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:14688-14706 Available from https://proceedings.mlr.press/v306/chen26aw.html.

Related Material