Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts

Hahyeon Choi, Nojun Kwak
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:20100-20126, 2026.

Abstract

We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and exhibits consistent sparsity-performance dynamics, exhibiting a reverse U-shaped trend, with performance peaking at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-choi26h, title = {Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts}, author = {Choi, Hahyeon and Kwak, Nojun}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {20100--20126}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/choi26h/choi26h.pdf}, url = {https://proceedings.mlr.press/v306/choi26h.html}, abstract = {We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and exhibits consistent sparsity-performance dynamics, exhibiting a reverse U-shaped trend, with performance peaking at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches.} }
Endnote
%0 Conference Paper %T Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts %A Hahyeon Choi %A Nojun Kwak %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-choi26h %I PMLR %P 20100--20126 %U https://proceedings.mlr.press/v306/choi26h.html %V 306 %X We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and exhibits consistent sparsity-performance dynamics, exhibiting a reverse U-shaped trend, with performance peaking at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches.
APA
Choi, H. & Kwak, N.. (2026). Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:20100-20126 Available from https://proceedings.mlr.press/v306/choi26h.html.

Related Material