[edit]
Constrained Random Forest for Domain-Generalizable Classification
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3888-3921, 2026.
Abstract
Domain generalization is a critical challenge in machine learning, where models must generalize outside their training domain(s) to unseen test or deployment domains without retraining. Motivated by the principle of independent causal mechanisms, the notion of domain invariance provides a principled causality-inspired approach to domain generalization. In this context, we introduce a general domain-constrained supervised learning framework where the aim is to learn a representation of the input data that results in improved domain invariance and stability in the resulting prediction model. Adapting the framework to tree-based learning, we introduce the Constrained Random Forest that prioritizes stable predictive mechanisms, leading to improved generalization across diverse environments. Experiments on synthetic and real-world datasets show improved robustness over competing methods, even in settings where the underlying assumptions that enable domain-invariant learning in theory are violated.