Determinantal Clustering Processes - A Nonparametric Bayesian Approach to Kernel Based Semi-Supervised Clustering

Amar Shah, Zoubin Ghahramani
Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence, PMLR R11:638-647, 2013.

Abstract

Semi-supervised clustering is the task of clus- tering data points into clusters where only a fraction of the points are labelled. The true number of clusters in the data is often un- known and most models require this param- eter as an input. Dirichlet process mixture models are appealing as they can infer the number of clusters from the data. However, these models do not deal with high dimen- sional data well and can encounter difficulties in inference. We present a novel nonparame- teric Bayesian method to cluster data points without the need to prespecify the number of clusters or to model complicated densities from which data points are assumed to be generated from. The key insight is to use determinants of submatrices of a kernel ma- trix as a measure of how close together a set of points are. We explore some theoretical properties of the model and derive a natural Gibbs based algorithm with MCMC hyper- parameter learning. We test the model on various synthetic and real world data sets.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR11-shah13a, title = {Determinantal Clustering Processes - A Nonparametric {B}ayesian Approach to Kernel Based Semi-Supervised Clustering}, author = {Shah, Amar and Ghahramani, Zoubin}, booktitle = {Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence}, pages = {638--647}, year = {2013}, editor = {Nicholson, Ann and Smyth, Padhraic}, volume = {R11}, series = {Proceedings of Machine Learning Research}, month = {12--14 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r11/main/assets/shah13a/shah13a.pdf}, url = {https://proceedings.mlr.press/r11/shah13a.html}, abstract = {Semi-supervised clustering is the task of clus- tering data points into clusters where only a fraction of the points are labelled. The true number of clusters in the data is often un- known and most models require this param- eter as an input. Dirichlet process mixture models are appealing as they can infer the number of clusters from the data. However, these models do not deal with high dimen- sional data well and can encounter difficulties in inference. We present a novel nonparame- teric Bayesian method to cluster data points without the need to prespecify the number of clusters or to model complicated densities from which data points are assumed to be generated from. The key insight is to use determinants of submatrices of a kernel ma- trix as a measure of how close together a set of points are. We explore some theoretical properties of the model and derive a natural Gibbs based algorithm with MCMC hyper- parameter learning. We test the model on various synthetic and real world data sets.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Determinantal Clustering Processes - A Nonparametric Bayesian Approach to Kernel Based Semi-Supervised Clustering %A Amar Shah %A Zoubin Ghahramani %B Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2013 %E Ann Nicholson %E Padhraic Smyth %F pmlr-vR11-shah13a %I PMLR %P 638--647 %U https://proceedings.mlr.press/r11/shah13a.html %V R11 %X Semi-supervised clustering is the task of clus- tering data points into clusters where only a fraction of the points are labelled. The true number of clusters in the data is often un- known and most models require this param- eter as an input. Dirichlet process mixture models are appealing as they can infer the number of clusters from the data. However, these models do not deal with high dimen- sional data well and can encounter difficulties in inference. We present a novel nonparame- teric Bayesian method to cluster data points without the need to prespecify the number of clusters or to model complicated densities from which data points are assumed to be generated from. The key insight is to use determinants of submatrices of a kernel ma- trix as a measure of how close together a set of points are. We explore some theoretical properties of the model and derive a natural Gibbs based algorithm with MCMC hyper- parameter learning. We test the model on various synthetic and real world data sets. %Z Reissued by PMLR on 04 October 2026.
APA
Shah, A. & Ghahramani, Z.. (2013). Determinantal Clustering Processes - A Nonparametric Bayesian Approach to Kernel Based Semi-Supervised Clustering. Proceedings of the 29th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R11:638-647 Available from https://proceedings.mlr.press/r11/shah13a.html. Reissued by PMLR on 04 October 2026.

Related Material