Causal Discovery for Efficient Offline RL with Factored Action Spaces

Cecilia Ehrlichman, Michael Dykstra, Shengpu Tang, Maggie Makar
Proceedings of the Fifth Conference on Causal Learning and Reasoning, PMLR 323:1424-1449, 2026.

Abstract

Offline policy optimization is often sample-inefficient, especially when the action space is large, a problem that commonly arises in healthcare applications and multi-agent tasks. Many domains, however, admit a combinatorial action space, where sub-actions affect future states and rewards independently of one another. Past work either makes a priori assumptions about sub-action independence leading to efficient but potentially biased policy optimization, or fails to leverage potential independence, sacrificing sample efficiency. In contrast, we propose a two-step framework that leverages causal discovery for efficient policy optimization without introducing bias. Our approach (i) discovers the causal structure underlying the environment’s dynamics from observational data, and (ii) exploits this structure to restrict the admissible policy class to a simpler, unbiased class. We provide theoretical guarantees characterizing settings under which our approach leads to efficient unbiased policy learning. Empirically, we demonstrate that our approach leads to more efficient policy optimization in settings with limited observational data, across both single-agent healthcare tasks and multi-agent settings. Our code is available at https://github.com/cehr123/DiFaRL.

Cite this Paper


BibTeX
@InProceedings{pmlr-v323-ehrlichman26a, title = {Causal Discovery for Efficient Offline RL with Factored Action Spaces}, author = {Ehrlichman, Cecilia and Dykstra, Michael and Tang, Shengpu and Makar, Maggie}, booktitle = {Proceedings of the Fifth Conference on Causal Learning and Reasoning}, pages = {1424--1449}, year = {2026}, editor = {Mazaheri, Bijan and Hanson, Niels Richard}, volume = {323}, series = {Proceedings of Machine Learning Research}, month = {06--08 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v323/main/assets/ehrlichman26a/ehrlichman26a.pdf}, url = {https://proceedings.mlr.press/v323/ehrlichman26a.html}, abstract = {Offline policy optimization is often sample-inefficient, especially when the action space is large, a problem that commonly arises in healthcare applications and multi-agent tasks. Many domains, however, admit a combinatorial action space, where sub-actions affect future states and rewards independently of one another. Past work either makes a priori assumptions about sub-action independence leading to efficient but potentially biased policy optimization, or fails to leverage potential independence, sacrificing sample efficiency. In contrast, we propose a two-step framework that leverages causal discovery for efficient policy optimization without introducing bias. Our approach (i) discovers the causal structure underlying the environment’s dynamics from observational data, and (ii) exploits this structure to restrict the admissible policy class to a simpler, unbiased class. We provide theoretical guarantees characterizing settings under which our approach leads to efficient unbiased policy learning. Empirically, we demonstrate that our approach leads to more efficient policy optimization in settings with limited observational data, across both single-agent healthcare tasks and multi-agent settings. Our code is available at https://github.com/cehr123/DiFaRL.} }
Endnote
%0 Conference Paper %T Causal Discovery for Efficient Offline RL with Factored Action Spaces %A Cecilia Ehrlichman %A Michael Dykstra %A Shengpu Tang %A Maggie Makar %B Proceedings of the Fifth Conference on Causal Learning and Reasoning %C Proceedings of Machine Learning Research %D 2026 %E Bijan Mazaheri %E Niels Richard Hanson %F pmlr-v323-ehrlichman26a %I PMLR %P 1424--1449 %U https://proceedings.mlr.press/v323/ehrlichman26a.html %V 323 %X Offline policy optimization is often sample-inefficient, especially when the action space is large, a problem that commonly arises in healthcare applications and multi-agent tasks. Many domains, however, admit a combinatorial action space, where sub-actions affect future states and rewards independently of one another. Past work either makes a priori assumptions about sub-action independence leading to efficient but potentially biased policy optimization, or fails to leverage potential independence, sacrificing sample efficiency. In contrast, we propose a two-step framework that leverages causal discovery for efficient policy optimization without introducing bias. Our approach (i) discovers the causal structure underlying the environment’s dynamics from observational data, and (ii) exploits this structure to restrict the admissible policy class to a simpler, unbiased class. We provide theoretical guarantees characterizing settings under which our approach leads to efficient unbiased policy learning. Empirically, we demonstrate that our approach leads to more efficient policy optimization in settings with limited observational data, across both single-agent healthcare tasks and multi-agent settings. Our code is available at https://github.com/cehr123/DiFaRL.
APA
Ehrlichman, C., Dykstra, M., Tang, S. & Makar, M.. (2026). Causal Discovery for Efficient Offline RL with Factored Action Spaces. Proceedings of the Fifth Conference on Causal Learning and Reasoning, in Proceedings of Machine Learning Research 323:1424-1449 Available from https://proceedings.mlr.press/v323/ehrlichman26a.html.

Related Material