Probabilistic Size-constrained Microclustering

Arto Klami, Aditya Jitta
Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, PMLR R14:445-454, 2016.

Abstract

Microclustering refers to clustering models that produce small clusters or, equivalently, to models where the size of the clusters grows sublinearly with the number of samples. We formulate probabilistic microclustering models by assigning a prior distribution on the size of the clusters, and in particular consider microclustering models with explicit bounds on the size of the clusters. The combinatorial constraints make full Bayesian inference complicated, but we manage to develop a Gibbs sampling algorithm that can efficiently sample from the joint cluster allocation of all data points. We empirically demonstrate the computational efficiency of the algorithm for problem instances of varying difficulty.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR14-klami16a, title = {Probabilistic Size-constrained Microclustering}, author = {Klami, Arto and Jitta, Aditya}, booktitle = {Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence}, pages = {445--454}, year = {2016}, editor = {Ihler, Alexander and Janzing, Dominik}, volume = {R14}, series = {Proceedings of Machine Learning Research}, month = {25--29 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r14/main/assets/klami16a/klami16a.pdf}, url = {https://proceedings.mlr.press/r14/klami16a.html}, abstract = {Microclustering refers to clustering models that produce small clusters or, equivalently, to models where the size of the clusters grows sublinearly with the number of samples. We formulate probabilistic microclustering models by assigning a prior distribution on the size of the clusters, and in particular consider microclustering models with explicit bounds on the size of the clusters. The combinatorial constraints make full Bayesian inference complicated, but we manage to develop a Gibbs sampling algorithm that can efficiently sample from the joint cluster allocation of all data points. We empirically demonstrate the computational efficiency of the algorithm for problem instances of varying difficulty.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Probabilistic Size-constrained Microclustering %A Arto Klami %A Aditya Jitta %B Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2016 %E Alexander Ihler %E Dominik Janzing %F pmlr-vR14-klami16a %I PMLR %P 445--454 %U https://proceedings.mlr.press/r14/klami16a.html %V R14 %X Microclustering refers to clustering models that produce small clusters or, equivalently, to models where the size of the clusters grows sublinearly with the number of samples. We formulate probabilistic microclustering models by assigning a prior distribution on the size of the clusters, and in particular consider microclustering models with explicit bounds on the size of the clusters. The combinatorial constraints make full Bayesian inference complicated, but we manage to develop a Gibbs sampling algorithm that can efficiently sample from the joint cluster allocation of all data points. We empirically demonstrate the computational efficiency of the algorithm for problem instances of varying difficulty. %Z Reissued by PMLR on 04 October 2026.
APA
Klami, A. & Jitta, A.. (2016). Probabilistic Size-constrained Microclustering. Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R14:445-454 Available from https://proceedings.mlr.press/r14/klami16a.html. Reissued by PMLR on 04 October 2026.

Related Material