From Token Imbalance to Balanced Routing: An ELBO-Regularized Probabilistic Framework for Contrastive Multimodal Learning

Habibeh Naderi, Behrouz Haji Soleimani, Stan Matwin
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2413-2421, 2026.

Abstract

We introduce CoPRIME (Contrastive Probabilistic Routing for IMbalanced tokens with ELBO-regularized mixture of experts), a probabilistic routing framework for multimodal representation learning that generalizes multimodal representation learning beyond vision-text by tackling the fundamental challenge of extreme token imbalance across modalities. This imbalanced-ness is particularly pronounced between spectrogram-tokenized audio and text. CoPRIME augments contrastive pretraining with an ELBO-regularized routing objective that jointly promotes 1) expert specialization, requiring experts to explain the tokens they receive, and 2) diverse utilization via KL regularization to a uniform prior. To stabilize routing, we further replace standard CoV-based regularizers with entropy-based importance and load losses, yielding smoother gradients and flexible, modality-aware routing without rigid uniformity constraints. On MOSEI and IEMOCAP datasets, CoPRIME achieves state-of-the-art zero- and few-shot emotion and sentiment results, outperforming dense Transformers and prior multimodal MoE variants while retaining the efficiency of sparse conditional computation. Ablations isolate the role of each loss and show that ELBO is the primary driver of stable specialization under modality imbalance, with entropy-based regularizers further improving convergence and utilization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-naderi26a, title = { From Token Imbalance to Balanced Routing: An ELBO-Regularized Probabilistic Framework for Contrastive Multimodal Learning }, author = {Naderi, Habibeh and Soleimani, Behrouz Haji and Matwin, Stan}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2413--2421}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/naderi26a/naderi26a.pdf}, url = {https://proceedings.mlr.press/v300/naderi26a.html}, abstract = { We introduce CoPRIME (Contrastive Probabilistic Routing for IMbalanced tokens with ELBO-regularized mixture of experts), a probabilistic routing framework for multimodal representation learning that generalizes multimodal representation learning beyond vision-text by tackling the fundamental challenge of extreme token imbalance across modalities. This imbalanced-ness is particularly pronounced between spectrogram-tokenized audio and text. CoPRIME augments contrastive pretraining with an ELBO-regularized routing objective that jointly promotes 1) expert specialization, requiring experts to explain the tokens they receive, and 2) diverse utilization via KL regularization to a uniform prior. To stabilize routing, we further replace standard CoV-based regularizers with entropy-based importance and load losses, yielding smoother gradients and flexible, modality-aware routing without rigid uniformity constraints. On MOSEI and IEMOCAP datasets, CoPRIME achieves state-of-the-art zero- and few-shot emotion and sentiment results, outperforming dense Transformers and prior multimodal MoE variants while retaining the efficiency of sparse conditional computation. Ablations isolate the role of each loss and show that ELBO is the primary driver of stable specialization under modality imbalance, with entropy-based regularizers further improving convergence and utilization. } }
Endnote
%0 Conference Paper %T From Token Imbalance to Balanced Routing: An ELBO-Regularized Probabilistic Framework for Contrastive Multimodal Learning %A Habibeh Naderi %A Behrouz Haji Soleimani %A Stan Matwin %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-naderi26a %I PMLR %P 2413--2421 %U https://proceedings.mlr.press/v300/naderi26a.html %V 300 %X We introduce CoPRIME (Contrastive Probabilistic Routing for IMbalanced tokens with ELBO-regularized mixture of experts), a probabilistic routing framework for multimodal representation learning that generalizes multimodal representation learning beyond vision-text by tackling the fundamental challenge of extreme token imbalance across modalities. This imbalanced-ness is particularly pronounced between spectrogram-tokenized audio and text. CoPRIME augments contrastive pretraining with an ELBO-regularized routing objective that jointly promotes 1) expert specialization, requiring experts to explain the tokens they receive, and 2) diverse utilization via KL regularization to a uniform prior. To stabilize routing, we further replace standard CoV-based regularizers with entropy-based importance and load losses, yielding smoother gradients and flexible, modality-aware routing without rigid uniformity constraints. On MOSEI and IEMOCAP datasets, CoPRIME achieves state-of-the-art zero- and few-shot emotion and sentiment results, outperforming dense Transformers and prior multimodal MoE variants while retaining the efficiency of sparse conditional computation. Ablations isolate the role of each loss and show that ELBO is the primary driver of stable specialization under modality imbalance, with entropy-based regularizers further improving convergence and utilization.
APA
Naderi, H., Soleimani, B.H. & Matwin, S.. (2026). From Token Imbalance to Balanced Routing: An ELBO-Regularized Probabilistic Framework for Contrastive Multimodal Learning . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2413-2421 Available from https://proceedings.mlr.press/v300/naderi26a.html.

Related Material