Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access

Daniel Ebi, Damien Ernst, Klemens Böhm, Gaspard Lambrechts
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27521-27549, 2026.

Abstract

Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability. Existing asymmetric actor-critic methods typically assume access to the full environment state to condition the critic during training, which is often unrealistic in practice. We introduce the informed asymmetric actor-critic framework that allows the critic to be conditioned on arbitrary state-dependent privileged signals, and show that any such signal yields unbiased policy gradient estimates. This substantially expands the set of admissible privileged information and raises the problem of selecting the most informative signals for learning. To this end, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a test based on improvements in value prediction that can be applied post hoc. Experiments on partially observable benchmarks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ebi26a, title = {Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access}, author = {Ebi, Daniel and Ernst, Damien and B\"{o}hm, Klemens and Lambrechts, Gaspard}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27521--27549}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ebi26a/ebi26a.pdf}, url = {https://proceedings.mlr.press/v306/ebi26a.html}, abstract = {Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability. Existing asymmetric actor-critic methods typically assume access to the full environment state to condition the critic during training, which is often unrealistic in practice. We introduce the informed asymmetric actor-critic framework that allows the critic to be conditioned on arbitrary state-dependent privileged signals, and show that any such signal yields unbiased policy gradient estimates. This substantially expands the set of admissible privileged information and raises the problem of selecting the most informative signals for learning. To this end, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a test based on improvements in value prediction that can be applied post hoc. Experiments on partially observable benchmarks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.} }
Endnote
%0 Conference Paper %T Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access %A Daniel Ebi %A Damien Ernst %A Klemens Böhm %A Gaspard Lambrechts %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ebi26a %I PMLR %P 27521--27549 %U https://proceedings.mlr.press/v306/ebi26a.html %V 306 %X Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability. Existing asymmetric actor-critic methods typically assume access to the full environment state to condition the critic during training, which is often unrealistic in practice. We introduce the informed asymmetric actor-critic framework that allows the critic to be conditioned on arbitrary state-dependent privileged signals, and show that any such signal yields unbiased policy gradient estimates. This substantially expands the set of admissible privileged information and raises the problem of selecting the most informative signals for learning. To this end, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a test based on improvements in value prediction that can be applied post hoc. Experiments on partially observable benchmarks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.
APA
Ebi, D., Ernst, D., Böhm, K. & Lambrechts, G.. (2026). Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27521-27549 Available from https://proceedings.mlr.press/v306/ebi26a.html.

Related Material