Causal Feature Learning via Generalized Rayleigh Quotients

Liang Cao, Jun Wan, Yan Qin, Weide Liu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:11496-11517, 2026.

Abstract

Extracting causally meaningful features from time-series data is fundamental for robust machine learning under distribution shifts. In process monitoring, existing methods struggle to maintain detection performance when operating conditions change. Current approaches capture either temporal causal relationships or cross-environment invariance, but not both simultaneously. We propose Causal Feature Learning (CFL), a unified framework that jointly optimizes for temporal relevance and environment mean invariance. CFL formulates feature extraction as a generalized Rayleigh-quotient problem, maximizing correlation with target variables while penalizing sensitivity to environment-dependent mean shifts. Theoretical analysis establishes conditions under which CFL identifies a mean-invariant predictive subspace. Experiments on the Tennessee Eastman Process demonstrate that CFL achieves 94.6% average fault detection rate, outperforming 15 baseline methods while operating at the lowest realized false-alarm rate and detection delay.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cao26t, title = {Causal Feature Learning via Generalized Rayleigh Quotients}, author = {Cao, Liang and Wan, Jun and Qin, Yan and Liu, Weide}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {11496--11517}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cao26t/cao26t.pdf}, url = {https://proceedings.mlr.press/v306/cao26t.html}, abstract = {Extracting causally meaningful features from time-series data is fundamental for robust machine learning under distribution shifts. In process monitoring, existing methods struggle to maintain detection performance when operating conditions change. Current approaches capture either temporal causal relationships or cross-environment invariance, but not both simultaneously. We propose Causal Feature Learning (CFL), a unified framework that jointly optimizes for temporal relevance and environment mean invariance. CFL formulates feature extraction as a generalized Rayleigh-quotient problem, maximizing correlation with target variables while penalizing sensitivity to environment-dependent mean shifts. Theoretical analysis establishes conditions under which CFL identifies a mean-invariant predictive subspace. Experiments on the Tennessee Eastman Process demonstrate that CFL achieves 94.6% average fault detection rate, outperforming 15 baseline methods while operating at the lowest realized false-alarm rate and detection delay.} }
Endnote
%0 Conference Paper %T Causal Feature Learning via Generalized Rayleigh Quotients %A Liang Cao %A Jun Wan %A Yan Qin %A Weide Liu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cao26t %I PMLR %P 11496--11517 %U https://proceedings.mlr.press/v306/cao26t.html %V 306 %X Extracting causally meaningful features from time-series data is fundamental for robust machine learning under distribution shifts. In process monitoring, existing methods struggle to maintain detection performance when operating conditions change. Current approaches capture either temporal causal relationships or cross-environment invariance, but not both simultaneously. We propose Causal Feature Learning (CFL), a unified framework that jointly optimizes for temporal relevance and environment mean invariance. CFL formulates feature extraction as a generalized Rayleigh-quotient problem, maximizing correlation with target variables while penalizing sensitivity to environment-dependent mean shifts. Theoretical analysis establishes conditions under which CFL identifies a mean-invariant predictive subspace. Experiments on the Tennessee Eastman Process demonstrate that CFL achieves 94.6% average fault detection rate, outperforming 15 baseline methods while operating at the lowest realized false-alarm rate and detection delay.
APA
Cao, L., Wan, J., Qin, Y. & Liu, W.. (2026). Causal Feature Learning via Generalized Rayleigh Quotients. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:11496-11517 Available from https://proceedings.mlr.press/v306/cao26t.html.

Related Material