Topological Alignment of Shared Vision-Language Embedding Space

Junwon You, Kang Dasol, Jae-Hun Jung
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1207-1215, 2026.

Abstract

Contrastive Vision-Language Models (VLMs) have demonstrated strong zero-shot capabilities. However, their cross-modal alignment remains biased toward English due to limited multilingual multimodal data. Recent multilingual extensions have alleviated this gap but enforce instance-level alignment while neglecting the global geometry of the shared embedding space. We address this problem by introducing \textbf{ToMCLIP} (\textbf{To}pological Alignment for \textbf{M}ultilingual \textbf{CLIP}), a topology-aware framework aligning embedding spaces with topology-preserving constraints. The proposed method applies persistent homology to define a topological alignment loss and approximates persistence diagram with theoretical error bounds using graph sparsification strategy. This work validates the proposed approach, showing enhanced structural coherence of multilingual representations, higher zero-shot accuracy on the CIFAR-100, and stronger multilingual retrieval performance on the xFlickr&CO. Beyond VLMs, the proposed approach provides a general method for incorporating topological alignment into representation learning. Code is available at \url{https://github.com/junwon0/ToMCLIP.git.}

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-you26a, title = { Topological Alignment of Shared Vision-Language Embedding Space }, author = {You, Junwon and Dasol, Kang and Jung, Jae-Hun}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1207--1215}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/you26a/you26a.pdf}, url = {https://proceedings.mlr.press/v300/you26a.html}, abstract = { Contrastive Vision-Language Models (VLMs) have demonstrated strong zero-shot capabilities. However, their cross-modal alignment remains biased toward English due to limited multilingual multimodal data. Recent multilingual extensions have alleviated this gap but enforce instance-level alignment while neglecting the global geometry of the shared embedding space. We address this problem by introducing \textbf{ToMCLIP} (\textbf{To}pological Alignment for \textbf{M}ultilingual \textbf{CLIP}), a topology-aware framework aligning embedding spaces with topology-preserving constraints. The proposed method applies persistent homology to define a topological alignment loss and approximates persistence diagram with theoretical error bounds using graph sparsification strategy. This work validates the proposed approach, showing enhanced structural coherence of multilingual representations, higher zero-shot accuracy on the CIFAR-100, and stronger multilingual retrieval performance on the xFlickr&CO. Beyond VLMs, the proposed approach provides a general method for incorporating topological alignment into representation learning. Code is available at \url{https://github.com/junwon0/ToMCLIP.git.} } }
Endnote
%0 Conference Paper %T Topological Alignment of Shared Vision-Language Embedding Space %A Junwon You %A Kang Dasol %A Jae-Hun Jung %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-you26a %I PMLR %P 1207--1215 %U https://proceedings.mlr.press/v300/you26a.html %V 300 %X Contrastive Vision-Language Models (VLMs) have demonstrated strong zero-shot capabilities. However, their cross-modal alignment remains biased toward English due to limited multilingual multimodal data. Recent multilingual extensions have alleviated this gap but enforce instance-level alignment while neglecting the global geometry of the shared embedding space. We address this problem by introducing \textbf{ToMCLIP} (\textbf{To}pological Alignment for \textbf{M}ultilingual \textbf{CLIP}), a topology-aware framework aligning embedding spaces with topology-preserving constraints. The proposed method applies persistent homology to define a topological alignment loss and approximates persistence diagram with theoretical error bounds using graph sparsification strategy. This work validates the proposed approach, showing enhanced structural coherence of multilingual representations, higher zero-shot accuracy on the CIFAR-100, and stronger multilingual retrieval performance on the xFlickr&CO. Beyond VLMs, the proposed approach provides a general method for incorporating topological alignment into representation learning. Code is available at \url{https://github.com/junwon0/ToMCLIP.git.}
APA
You, J., Dasol, K. & Jung, J.. (2026). Topological Alignment of Shared Vision-Language Embedding Space . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1207-1215 Available from https://proceedings.mlr.press/v300/you26a.html.

Related Material