Consistent PCA and Spectral Clustering

Satoshi Hara, Yuichi Yoshida
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:910-918, 2026.

Abstract

Principal component analysis (PCA) and spectral clustering are representative methods for extracting and interpreting the inherent structure of data. However, if the output results significantly change upon the addition of new data points, it can lead to several issues such as instability in the downstream task or a lack of trust in the findings. To address these problems, we consider online variants of PCA and spectral clustering, and show that a natural subspace-preserving regularizer provides provable approximation and consistency guarantees. Here, an algorithm is said to have a high consistency if the output change, with respect to an appropriate distance metric, is small when new data points are added. We empirically confirm the superiority of the proposed methods using real-world data.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-hara26a, title = { Consistent PCA and Spectral Clustering }, author = {Hara, Satoshi and Yoshida, Yuichi}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {910--918}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/hara26a/hara26a.pdf}, url = {https://proceedings.mlr.press/v300/hara26a.html}, abstract = { Principal component analysis (PCA) and spectral clustering are representative methods for extracting and interpreting the inherent structure of data. However, if the output results significantly change upon the addition of new data points, it can lead to several issues such as instability in the downstream task or a lack of trust in the findings. To address these problems, we consider online variants of PCA and spectral clustering, and show that a natural subspace-preserving regularizer provides provable approximation and consistency guarantees. Here, an algorithm is said to have a high consistency if the output change, with respect to an appropriate distance metric, is small when new data points are added. We empirically confirm the superiority of the proposed methods using real-world data. } }
Endnote
%0 Conference Paper %T Consistent PCA and Spectral Clustering %A Satoshi Hara %A Yuichi Yoshida %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-hara26a %I PMLR %P 910--918 %U https://proceedings.mlr.press/v300/hara26a.html %V 300 %X Principal component analysis (PCA) and spectral clustering are representative methods for extracting and interpreting the inherent structure of data. However, if the output results significantly change upon the addition of new data points, it can lead to several issues such as instability in the downstream task or a lack of trust in the findings. To address these problems, we consider online variants of PCA and spectral clustering, and show that a natural subspace-preserving regularizer provides provable approximation and consistency guarantees. Here, an algorithm is said to have a high consistency if the output change, with respect to an appropriate distance metric, is small when new data points are added. We empirically confirm the superiority of the proposed methods using real-world data.
APA
Hara, S. & Yoshida, Y.. (2026). Consistent PCA and Spectral Clustering . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:910-918 Available from https://proceedings.mlr.press/v300/hara26a.html.

Related Material