Transformation Based Probabilistic Clustering using Supervision

Siddharth Gopal Carnegie Mellon University, Yiming Yang Carnegie Mellon University
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:304-313, 2014.

Abstract

One of the common problems with clustering is that the generated clusters often do not match user expectations. This paper proposes a novel probabilistic framework that exploits supervised information in a discriminative and transferable manner to generate better clustering of unlabeled data. The supervision is provided by revealing the cluster assignments for some subset of the ground truth clusters and is used to learn a trans- formation of the data such that labeled instances form well-separated clusters with respect to the given clustering objective. This estimated trans- formation function enables us to fold the remain- ing unlabeled data into a space where new clus- ters hopefully match user expectations. While our framework is general, in this paper, we fo- cus on its application to Gaussian and von Mises- Fisher mixture models. Extensive testing on 23 data sets across several application domains re- vealed substantial improvement in performance over competing methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-university14j, title = {Transformation Based Probabilistic Clustering using Supervision}, author = {University, Siddharth Gopal Carnegie Mellon and University, Yiming Yang Carnegie Mellon}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {304--313}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/university14j/university14j.pdf}, url = {https://proceedings.mlr.press/r12/university14j.html}, abstract = {One of the common problems with clustering is that the generated clusters often do not match user expectations. This paper proposes a novel probabilistic framework that exploits supervised information in a discriminative and transferable manner to generate better clustering of unlabeled data. The supervision is provided by revealing the cluster assignments for some subset of the ground truth clusters and is used to learn a trans- formation of the data such that labeled instances form well-separated clusters with respect to the given clustering objective. This estimated trans- formation function enables us to fold the remain- ing unlabeled data into a space where new clus- ters hopefully match user expectations. While our framework is general, in this paper, we fo- cus on its application to Gaussian and von Mises- Fisher mixture models. Extensive testing on 23 data sets across several application domains re- vealed substantial improvement in performance over competing methods.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Transformation Based Probabilistic Clustering using Supervision %A Siddharth Gopal Carnegie Mellon University %A Yiming Yang Carnegie Mellon University %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-university14j %I PMLR %P 304--313 %U https://proceedings.mlr.press/r12/university14j.html %V R12 %X One of the common problems with clustering is that the generated clusters often do not match user expectations. This paper proposes a novel probabilistic framework that exploits supervised information in a discriminative and transferable manner to generate better clustering of unlabeled data. The supervision is provided by revealing the cluster assignments for some subset of the ground truth clusters and is used to learn a trans- formation of the data such that labeled instances form well-separated clusters with respect to the given clustering objective. This estimated trans- formation function enables us to fold the remain- ing unlabeled data into a space where new clus- ters hopefully match user expectations. While our framework is general, in this paper, we fo- cus on its application to Gaussian and von Mises- Fisher mixture models. Extensive testing on 23 data sets across several application domains re- vealed substantial improvement in performance over competing methods. %Z Reissued by PMLR on 04 October 2026.
APA
University, S.G.C.M. & University, Y.Y.C.M.. (2014). Transformation Based Probabilistic Clustering using Supervision. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:304-313 Available from https://proceedings.mlr.press/r12/university14j.html. Reissued by PMLR on 04 October 2026.

Related Material