Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms

Katherine Avery, Chinmay Pendse, David Jensen
Proceedings of the Fifth Conference on Causal Learning and Reasoning, PMLR 323:593-626, 2026.

Abstract

Causal graphical models can encode large amounts structural knowledge, both from the background knowledge of domain experts and the structural knowledge discovered from randomized experiments or observational data. However, though we may know the general structure of causal relationships, we often do not know the exact causal mechanisms. In this work, we propose a causal multi-armed bandit evaluation and learning algorithm that can reason effectively despite uncertainty over conditional probability distributions. Further, we show how conditional independence testing can be used to choose variables for modeling. We find that the structural equation model (SEM) approach gives more accurate evaluations compared to traditional approaches, particularly as the range of possible causal mechanisms grows. Further, the SEM approach learns low-variance policies, and it learns an optimal policy, assuming the model is sufficiently well-specified. Traditional approaches can converge to local extrema or fail to converge at all.

Cite this Paper


BibTeX
@InProceedings{pmlr-v323-avery26a, title = {Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms}, author = {Avery, Katherine and Pendse, Chinmay and Jensen, David}, booktitle = {Proceedings of the Fifth Conference on Causal Learning and Reasoning}, pages = {593--626}, year = {2026}, editor = {Mazaheri, Bijan and Hanson, Niels Richard}, volume = {323}, series = {Proceedings of Machine Learning Research}, month = {06--08 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v323/main/assets/avery26a/avery26a.pdf}, url = {https://proceedings.mlr.press/v323/avery26a.html}, abstract = {Causal graphical models can encode large amounts structural knowledge, both from the background knowledge of domain experts and the structural knowledge discovered from randomized experiments or observational data. However, though we may know the general structure of causal relationships, we often do not know the exact causal mechanisms. In this work, we propose a causal multi-armed bandit evaluation and learning algorithm that can reason effectively despite uncertainty over conditional probability distributions. Further, we show how conditional independence testing can be used to choose variables for modeling. We find that the structural equation model (SEM) approach gives more accurate evaluations compared to traditional approaches, particularly as the range of possible causal mechanisms grows. Further, the SEM approach learns low-variance policies, and it learns an optimal policy, assuming the model is sufficiently well-specified. Traditional approaches can converge to local extrema or fail to converge at all.} }
Endnote
%0 Conference Paper %T Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms %A Katherine Avery %A Chinmay Pendse %A David Jensen %B Proceedings of the Fifth Conference on Causal Learning and Reasoning %C Proceedings of Machine Learning Research %D 2026 %E Bijan Mazaheri %E Niels Richard Hanson %F pmlr-v323-avery26a %I PMLR %P 593--626 %U https://proceedings.mlr.press/v323/avery26a.html %V 323 %X Causal graphical models can encode large amounts structural knowledge, both from the background knowledge of domain experts and the structural knowledge discovered from randomized experiments or observational data. However, though we may know the general structure of causal relationships, we often do not know the exact causal mechanisms. In this work, we propose a causal multi-armed bandit evaluation and learning algorithm that can reason effectively despite uncertainty over conditional probability distributions. Further, we show how conditional independence testing can be used to choose variables for modeling. We find that the structural equation model (SEM) approach gives more accurate evaluations compared to traditional approaches, particularly as the range of possible causal mechanisms grows. Further, the SEM approach learns low-variance policies, and it learns an optimal policy, assuming the model is sufficiently well-specified. Traditional approaches can converge to local extrema or fail to converge at all.
APA
Avery, K., Pendse, C. & Jensen, D.. (2026). Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms. Proceedings of the Fifth Conference on Causal Learning and Reasoning, in Proceedings of Machine Learning Research 323:593-626 Available from https://proceedings.mlr.press/v323/avery26a.html.

Related Material