A Causal Decomposition Approach for Fair Contextual Multi-Armed Bandits

Jiajun Chen, Jin Tian, Christopher John Quinn
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16945-16995, 2026.

Abstract

Counterfactual reasoning is one of the fundamental facets of human cognition, involved in various tasks such as explanation, credit assignment, blame, and responsibility. It describes the queries what would have happened had some intervention been performed given that something else, corresponding to Layer 3 of the Pearl Causal Hierarchy. In this project, we examine specific types of counterfactual quantities, called counterfactual direct ($\mathrm{Ctf}\text{-}\mathrm{DE}$), indirect ($\mathrm{Ctf}\text{-}\mathrm{IE}$), and spurious ($\mathrm{Ctf}\text{-}\mathrm{SE}$) effects for quantifying fairness in a sequential decision-making framework. Building on these measures, we formulate an online causally-fair learning problem with multiple long-term constraints and study it in both non-parametric contextual bandits and parametric logistic bandits settings. We achieve sublinear regret and violations bounds for both bandits settings with roundwise counterfactual fairness constraints (that are a priori unknown) without Slater’s condition. For logistic bandits, our method achieves a regret with leading term matches that of the unconstrained setting.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26ef, title = {A Causal Decomposition Approach for Fair Contextual Multi-Armed Bandits}, author = {Chen, Jiajun and Tian, Jin and Quinn, Christopher John}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16945--16995}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26ef/chen26ef.pdf}, url = {https://proceedings.mlr.press/v306/chen26ef.html}, abstract = {Counterfactual reasoning is one of the fundamental facets of human cognition, involved in various tasks such as explanation, credit assignment, blame, and responsibility. It describes the queries what would have happened had some intervention been performed given that something else, corresponding to Layer 3 of the Pearl Causal Hierarchy. In this project, we examine specific types of counterfactual quantities, called counterfactual direct ($\mathrm{Ctf}\text{-}\mathrm{DE}$), indirect ($\mathrm{Ctf}\text{-}\mathrm{IE}$), and spurious ($\mathrm{Ctf}\text{-}\mathrm{SE}$) effects for quantifying fairness in a sequential decision-making framework. Building on these measures, we formulate an online causally-fair learning problem with multiple long-term constraints and study it in both non-parametric contextual bandits and parametric logistic bandits settings. We achieve sublinear regret and violations bounds for both bandits settings with roundwise counterfactual fairness constraints (that are a priori unknown) without Slater’s condition. For logistic bandits, our method achieves a regret with leading term matches that of the unconstrained setting.} }
Endnote
%0 Conference Paper %T A Causal Decomposition Approach for Fair Contextual Multi-Armed Bandits %A Jiajun Chen %A Jin Tian %A Christopher John Quinn %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26ef %I PMLR %P 16945--16995 %U https://proceedings.mlr.press/v306/chen26ef.html %V 306 %X Counterfactual reasoning is one of the fundamental facets of human cognition, involved in various tasks such as explanation, credit assignment, blame, and responsibility. It describes the queries what would have happened had some intervention been performed given that something else, corresponding to Layer 3 of the Pearl Causal Hierarchy. In this project, we examine specific types of counterfactual quantities, called counterfactual direct ($\mathrm{Ctf}\text{-}\mathrm{DE}$), indirect ($\mathrm{Ctf}\text{-}\mathrm{IE}$), and spurious ($\mathrm{Ctf}\text{-}\mathrm{SE}$) effects for quantifying fairness in a sequential decision-making framework. Building on these measures, we formulate an online causally-fair learning problem with multiple long-term constraints and study it in both non-parametric contextual bandits and parametric logistic bandits settings. We achieve sublinear regret and violations bounds for both bandits settings with roundwise counterfactual fairness constraints (that are a priori unknown) without Slater’s condition. For logistic bandits, our method achieves a regret with leading term matches that of the unconstrained setting.
APA
Chen, J., Tian, J. & Quinn, C.J.. (2026). A Causal Decomposition Approach for Fair Contextual Multi-Armed Bandits. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16945-16995 Available from https://proceedings.mlr.press/v306/chen26ef.html.

Related Material