[edit]
Beyond Bounds: Quantifying the Probability of Counterfactual Fairness
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2994-3011, 2026.
Abstract
Counterfactual fairness is a rigorous criterion for algorithmic decision-making but remains fundamentally unidentifiable from observational data alone. Existing partial identification methods address this by deriving bounds for fairness measures; however, these intervals are often wide, limiting their practical utility. To address this limitation, we propose a framework that quantifies the probability that a black-box model satisfies counterfactual fairness, under a stated prior over the structural causal models compatible with the observed data. By exploiting conditional independencies to reduce the parameter space, we demonstrate that the set of causal parameters compatible with the observed data forms a convex polytope. We further show that, given domain-specific priors on exogenous distributions, this prior-dependent probability can be estimated via Hit-and-Run Monte Carlo integration. Our approach provides a practical tool that complements worst-case bounds, offering a prior-dependent probabilistic summary for auditing model fairness.