Constrained Random Forest for Domain-Generalizable Classification

Camilla Lingjærde, Geir Kjetil Sandve, Arnoldo Frigessi, Sylvia Richardson, Johan Pensar
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3888-3921, 2026.

Abstract

Domain generalization is a critical challenge in machine learning, where models must generalize outside their training domain(s) to unseen test or deployment domains without retraining. Motivated by the principle of independent causal mechanisms, the notion of domain invariance provides a principled causality-inspired approach to domain generalization. In this context, we introduce a general domain-constrained supervised learning framework where the aim is to learn a representation of the input data that results in improved domain invariance and stability in the resulting prediction model. Adapting the framework to tree-based learning, we introduce the Constrained Random Forest that prioritizes stable predictive mechanisms, leading to improved generalization across diverse environments. Experiments on synthetic and real-world datasets show improved robustness over competing methods, even in settings where the underlying assumptions that enable domain-invariant learning in theory are violated.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-lingjaerde26a, title = {Constrained Random Forest for Domain-Generalizable Classification}, author = {Lingj\ae{}rde, Camilla and Sandve, Geir Kjetil and Frigessi, Arnoldo and Richardson, Sylvia and Pensar, Johan}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3888--3921}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/lingjaerde26a/lingjaerde26a.pdf}, url = {https://proceedings.mlr.press/v337/lingjaerde26a.html}, abstract = {Domain generalization is a critical challenge in machine learning, where models must generalize outside their training domain(s) to unseen test or deployment domains without retraining. Motivated by the principle of independent causal mechanisms, the notion of domain invariance provides a principled causality-inspired approach to domain generalization. In this context, we introduce a general domain-constrained supervised learning framework where the aim is to learn a representation of the input data that results in improved domain invariance and stability in the resulting prediction model. Adapting the framework to tree-based learning, we introduce the Constrained Random Forest that prioritizes stable predictive mechanisms, leading to improved generalization across diverse environments. Experiments on synthetic and real-world datasets show improved robustness over competing methods, even in settings where the underlying assumptions that enable domain-invariant learning in theory are violated.} }
Endnote
%0 Conference Paper %T Constrained Random Forest for Domain-Generalizable Classification %A Camilla Lingjærde %A Geir Kjetil Sandve %A Arnoldo Frigessi %A Sylvia Richardson %A Johan Pensar %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-lingjaerde26a %I PMLR %P 3888--3921 %U https://proceedings.mlr.press/v337/lingjaerde26a.html %V 337 %X Domain generalization is a critical challenge in machine learning, where models must generalize outside their training domain(s) to unseen test or deployment domains without retraining. Motivated by the principle of independent causal mechanisms, the notion of domain invariance provides a principled causality-inspired approach to domain generalization. In this context, we introduce a general domain-constrained supervised learning framework where the aim is to learn a representation of the input data that results in improved domain invariance and stability in the resulting prediction model. Adapting the framework to tree-based learning, we introduce the Constrained Random Forest that prioritizes stable predictive mechanisms, leading to improved generalization across diverse environments. Experiments on synthetic and real-world datasets show improved robustness over competing methods, even in settings where the underlying assumptions that enable domain-invariant learning in theory are violated.
APA
Lingjærde, C., Sandve, G.K., Frigessi, A., Richardson, S. & Pensar, J.. (2026). Constrained Random Forest for Domain-Generalizable Classification. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3888-3921 Available from https://proceedings.mlr.press/v337/lingjaerde26a.html.

Related Material