[edit]
Provable Subspace Identification of Nonlinear Multi-view CCA
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1947-1982, 2026.
Abstract
We investigate the identifiability of nonlinear canonical correlation analysis ({CCA}) in a multi-view setup, in which each view is generated by applying an unknown nonlinear map to a linear mixture of shared latent variables plus view-private noise. Rather than pursuing exact unmixing, which is known to be ill-posed under general nonlinear mixing, we instead reframe multi-view {CCA} as a basis-invariant subspace identification problem. Under suitable latent priors and spectral separation conditions, we prove that the pairwise population {CCA} objective recovers correlated signal subspaces up to view-wise orthogonal ambiguity. For $N \geq 3$ views, their multi-view aggregation provably isolates the jointly correlated subspaces shared across all views while eliminating view-private variation. We further establish finite-sample statistical consistency guarantees by translating the concentration of empirical cross-covariances into explicit subspace error bounds via spectral perturbation theory. Experiments on synthetic and rendered image datasets support our theoretical findings and illustrate the necessity of the assumed conditions.