Multilingual Topic Models for Unaligned Text

Jordan Boyd-Graber, David Blei
Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, PMLR R7:67-74, 2009.

Abstract

We develop the multilingual topic model for unaligned text (MuTo), a probabilistic model of text that is designed to analyze corpora composed of documents in two languages. From these documents, MuTo uses stochastic EM to simultaneously discover both a matching between the languages and multilingual latent topics. We demonstrate that MuTo is able to find shared topics on real-world multilingual corpora, successfully pairing related documents across languages. MuTo provides a new framework for creating multilingual topic models without needing carefully curated parallel corpora and allows applications built using the topic model formalism to be applied to a much wider class of corpora.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR7-boyd-graber09a, title = {Multilingual Topic Models for Unaligned Text}, author = {Boyd-Graber, Jordan and Blei, David}, booktitle = {Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence}, pages = {67--74}, year = {2009}, editor = {Bilmes, Jeff and Ng, Andrew Y.}, volume = {R7}, series = {Proceedings of Machine Learning Research}, month = {18--21 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r7/main/assets/boyd-graber09a/boyd-graber09a.pdf}, url = {https://proceedings.mlr.press/r7/boyd-graber09a.html}, abstract = {We develop the multilingual topic model for unaligned text (MuTo), a probabilistic model of text that is designed to analyze corpora composed of documents in two languages. From these documents, MuTo uses stochastic EM to simultaneously discover both a matching between the languages and multilingual latent topics. We demonstrate that MuTo is able to find shared topics on real-world multilingual corpora, successfully pairing related documents across languages. MuTo provides a new framework for creating multilingual topic models without needing carefully curated parallel corpora and allows applications built using the topic model formalism to be applied to a much wider class of corpora.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Multilingual Topic Models for Unaligned Text %A Jordan Boyd-Graber %A David Blei %B Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2009 %E Jeff Bilmes %E Andrew Y. Ng %F pmlr-vR7-boyd-graber09a %I PMLR %P 67--74 %U https://proceedings.mlr.press/r7/boyd-graber09a.html %V R7 %X We develop the multilingual topic model for unaligned text (MuTo), a probabilistic model of text that is designed to analyze corpora composed of documents in two languages. From these documents, MuTo uses stochastic EM to simultaneously discover both a matching between the languages and multilingual latent topics. We demonstrate that MuTo is able to find shared topics on real-world multilingual corpora, successfully pairing related documents across languages. MuTo provides a new framework for creating multilingual topic models without needing carefully curated parallel corpora and allows applications built using the topic model formalism to be applied to a much wider class of corpora. %Z Reissued by PMLR on 04 October 2026.
APA
Boyd-Graber, J. & Blei, D.. (2009). Multilingual Topic Models for Unaligned Text. Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R7:67-74 Available from https://proceedings.mlr.press/r7/boyd-graber09a.html. Reissued by PMLR on 04 October 2026.

Related Material