A Correlation Analysis Approach to Finding Interpretable Latent Representations via Conditional Generative Models

James Buenfil, Eardi Lila
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4816-4824, 2026.

Abstract

Supervised disentanglement, that is, learning interpretable nonlinear latent representations of a target data view informed by an auxiliary data view, is a central challenge in interpretable machine learning. We formulate this problem as a partially linear invertible canonical correlation analysis (PLiCCA). Specifically, given two data views, (i) complex data lying near a potentially high-dimensional manifold, and (ii) auxiliary high-dimensional multivariate data, PLiCCA learns latent variables for the complex view that are maximally correlated with sparse linear combinations of the auxiliary variables. In contrast to regression-based approaches to supervised disentanglement, the proposed method yields a latent embedding whose coordinates are explicitly ordered by their interpretability with respect to the auxiliary variables. We formalize the population PLiCCA problem and establish existence results. We then show a close theoretical connection between PLiCCA and conditional latent variable models, in particular conditional variational autoencoders and conditional normalizing flows, which enables practical estimation. We demonstrate our approach on brain imaging data, where PLiCCA is used to learn embeddings informed by auxiliary demographic, psychometric, and behavioral variables.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-buenfil26a, title = { A Correlation Analysis Approach to Finding Interpretable Latent Representations via Conditional Generative Models }, author = {Buenfil, James and Lila, Eardi}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4816--4824}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/buenfil26a/buenfil26a.pdf}, url = {https://proceedings.mlr.press/v300/buenfil26a.html}, abstract = { Supervised disentanglement, that is, learning interpretable nonlinear latent representations of a target data view informed by an auxiliary data view, is a central challenge in interpretable machine learning. We formulate this problem as a partially linear invertible canonical correlation analysis (PLiCCA). Specifically, given two data views, (i) complex data lying near a potentially high-dimensional manifold, and (ii) auxiliary high-dimensional multivariate data, PLiCCA learns latent variables for the complex view that are maximally correlated with sparse linear combinations of the auxiliary variables. In contrast to regression-based approaches to supervised disentanglement, the proposed method yields a latent embedding whose coordinates are explicitly ordered by their interpretability with respect to the auxiliary variables. We formalize the population PLiCCA problem and establish existence results. We then show a close theoretical connection between PLiCCA and conditional latent variable models, in particular conditional variational autoencoders and conditional normalizing flows, which enables practical estimation. We demonstrate our approach on brain imaging data, where PLiCCA is used to learn embeddings informed by auxiliary demographic, psychometric, and behavioral variables. } }
Endnote
%0 Conference Paper %T A Correlation Analysis Approach to Finding Interpretable Latent Representations via Conditional Generative Models %A James Buenfil %A Eardi Lila %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-buenfil26a %I PMLR %P 4816--4824 %U https://proceedings.mlr.press/v300/buenfil26a.html %V 300 %X Supervised disentanglement, that is, learning interpretable nonlinear latent representations of a target data view informed by an auxiliary data view, is a central challenge in interpretable machine learning. We formulate this problem as a partially linear invertible canonical correlation analysis (PLiCCA). Specifically, given two data views, (i) complex data lying near a potentially high-dimensional manifold, and (ii) auxiliary high-dimensional multivariate data, PLiCCA learns latent variables for the complex view that are maximally correlated with sparse linear combinations of the auxiliary variables. In contrast to regression-based approaches to supervised disentanglement, the proposed method yields a latent embedding whose coordinates are explicitly ordered by their interpretability with respect to the auxiliary variables. We formalize the population PLiCCA problem and establish existence results. We then show a close theoretical connection between PLiCCA and conditional latent variable models, in particular conditional variational autoencoders and conditional normalizing flows, which enables practical estimation. We demonstrate our approach on brain imaging data, where PLiCCA is used to learn embeddings informed by auxiliary demographic, psychometric, and behavioral variables.
APA
Buenfil, J. & Lila, E.. (2026). A Correlation Analysis Approach to Finding Interpretable Latent Representations via Conditional Generative Models . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4816-4824 Available from https://proceedings.mlr.press/v300/buenfil26a.html.

Related Material