[edit]
A Causal Decomposition Approach for Fair Contextual Multi-Armed Bandits
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16945-16995, 2026.
Abstract
Counterfactual reasoning is one of the fundamental facets of human cognition, involved in various tasks such as explanation, credit assignment, blame, and responsibility. It describes the queries what would have happened had some intervention been performed given that something else, corresponding to Layer 3 of the Pearl Causal Hierarchy. In this project, we examine specific types of counterfactual quantities, called counterfactual direct ($\mathrm{Ctf}\text{-}\mathrm{DE}$), indirect ($\mathrm{Ctf}\text{-}\mathrm{IE}$), and spurious ($\mathrm{Ctf}\text{-}\mathrm{SE}$) effects for quantifying fairness in a sequential decision-making framework. Building on these measures, we formulate an online causally-fair learning problem with multiple long-term constraints and study it in both non-parametric contextual bandits and parametric logistic bandits settings. We achieve sublinear regret and violations bounds for both bandits settings with roundwise counterfactual fairness constraints (that are a priori unknown) without Slater’s condition. For logistic bandits, our method achieves a regret with leading term matches that of the unconstrained setting.