Scalable Nonparametric Bayesian Multilevel Clustering

Viet Huynh, Dinh Phung, Svetha Venkatesh, XuanLong Nguyen, Matthew Hoffman, Hung Bui
Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, PMLR R14:732-741, 2016.

Abstract

Multilevel clustering problems where the content and contextual information are jointly clustered are ubiquitous in modern data sets. Existing work on this problem are limited to small datasets due to the use of the Gibbs sampler. We address the problem of scaling up multilevel clustering under a Bayesian nonparametric setting, extending the MC2 model proposed in (Nguyen et al., 2014). We ground our approach in mean-field and stochastic variational inference (SVI) theory. However, the interplay between content and context modeling makes naive mean-field approach inefficient. We develop a tree-structured SVI algorithm that avoids the need to repeatedly go through the corpus as in Gibbs sampler. More crucially, our method is immediately amendable to parallelization, facilitating a scalable distributed implementation of our algorithm on the Apache Spark platform. We conducted extensive experiments in a variety of domains including text, images, and real-world user application activities. Direct comparison with the Gibbs-sampler demonstrates that our method is an order-of-magnitude faster without loss of model quality. Our Spark-based implementation gains another order-of-magnitude speed up and can scale to large real-world data sets containing millions of documents and groups.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR14-huynh16a, title = {Scalable Nonparametric {B}ayesian Multilevel Clustering}, author = {Huynh, Viet and Phung, Dinh and Venkatesh, Svetha and Nguyen, XuanLong and Hoffman, Matthew and Bui, Hung}, booktitle = {Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence}, pages = {732--741}, year = {2016}, editor = {Ihler, Alexander and Janzing, Dominik}, volume = {R14}, series = {Proceedings of Machine Learning Research}, month = {25--29 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r14/main/assets/huynh16a/huynh16a.pdf}, url = {https://proceedings.mlr.press/r14/huynh16a.html}, abstract = {Multilevel clustering problems where the content and contextual information are jointly clustered are ubiquitous in modern data sets. Existing work on this problem are limited to small datasets due to the use of the Gibbs sampler. We address the problem of scaling up multilevel clustering under a Bayesian nonparametric setting, extending the MC2 model proposed in (Nguyen et al., 2014). We ground our approach in mean-field and stochastic variational inference (SVI) theory. However, the interplay between content and context modeling makes naive mean-field approach inefficient. We develop a tree-structured SVI algorithm that avoids the need to repeatedly go through the corpus as in Gibbs sampler. More crucially, our method is immediately amendable to parallelization, facilitating a scalable distributed implementation of our algorithm on the Apache Spark platform. We conducted extensive experiments in a variety of domains including text, images, and real-world user application activities. Direct comparison with the Gibbs-sampler demonstrates that our method is an order-of-magnitude faster without loss of model quality. Our Spark-based implementation gains another order-of-magnitude speed up and can scale to large real-world data sets containing millions of documents and groups.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Scalable Nonparametric Bayesian Multilevel Clustering %A Viet Huynh %A Dinh Phung %A Svetha Venkatesh %A XuanLong Nguyen %A Matthew Hoffman %A Hung Bui %B Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2016 %E Alexander Ihler %E Dominik Janzing %F pmlr-vR14-huynh16a %I PMLR %P 732--741 %U https://proceedings.mlr.press/r14/huynh16a.html %V R14 %X Multilevel clustering problems where the content and contextual information are jointly clustered are ubiquitous in modern data sets. Existing work on this problem are limited to small datasets due to the use of the Gibbs sampler. We address the problem of scaling up multilevel clustering under a Bayesian nonparametric setting, extending the MC2 model proposed in (Nguyen et al., 2014). We ground our approach in mean-field and stochastic variational inference (SVI) theory. However, the interplay between content and context modeling makes naive mean-field approach inefficient. We develop a tree-structured SVI algorithm that avoids the need to repeatedly go through the corpus as in Gibbs sampler. More crucially, our method is immediately amendable to parallelization, facilitating a scalable distributed implementation of our algorithm on the Apache Spark platform. We conducted extensive experiments in a variety of domains including text, images, and real-world user application activities. Direct comparison with the Gibbs-sampler demonstrates that our method is an order-of-magnitude faster without loss of model quality. Our Spark-based implementation gains another order-of-magnitude speed up and can scale to large real-world data sets containing millions of documents and groups. %Z Reissued by PMLR on 04 October 2026.
APA
Huynh, V., Phung, D., Venkatesh, S., Nguyen, X., Hoffman, M. & Bui, H.. (2016). Scalable Nonparametric Bayesian Multilevel Clustering. Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R14:732-741 Available from https://proceedings.mlr.press/r14/huynh16a.html. Reissued by PMLR on 04 October 2026.

Related Material