Hybrid Reinforcement Learning in Adversarial Markov Decision Processes

Duo Cheng, Xingyu Zhou, Bo Ji
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:19179-19210, 2026.

Abstract

We study hybrid reinforcement learning (RL) in adversarial Markov Decision Processes (MDPs), where the learner simultaneously receives on-policy feedback from the executed policy and off-policy feedback from a fixed behavior policy, and loss functions can change arbitrarily over time. On-policy feedback allows exploration and ensures the worst-case guarantee against any comparator policy, while off-policy feedback provides coverage-dependent guarantee that scales with the "mismatch" between the behavior and comparator policies (called coverage ratio) and can be sharper than on-policy results whenever this ratio is small. We propose a new hybrid RL framework that accommodates adversarial losses and unknown transitions, preserving off-policy guarantees while ensuring non-trivial worst-case performance.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cheng26s, title = {Hybrid Reinforcement Learning in Adversarial {M}arkov Decision Processes}, author = {Cheng, Duo and Zhou, Xingyu and Ji, Bo}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {19179--19210}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cheng26s/cheng26s.pdf}, url = {https://proceedings.mlr.press/v306/cheng26s.html}, abstract = {We study hybrid reinforcement learning (RL) in adversarial Markov Decision Processes (MDPs), where the learner simultaneously receives on-policy feedback from the executed policy and off-policy feedback from a fixed behavior policy, and loss functions can change arbitrarily over time. On-policy feedback allows exploration and ensures the worst-case guarantee against any comparator policy, while off-policy feedback provides coverage-dependent guarantee that scales with the "mismatch" between the behavior and comparator policies (called coverage ratio) and can be sharper than on-policy results whenever this ratio is small. We propose a new hybrid RL framework that accommodates adversarial losses and unknown transitions, preserving off-policy guarantees while ensuring non-trivial worst-case performance.} }
Endnote
%0 Conference Paper %T Hybrid Reinforcement Learning in Adversarial Markov Decision Processes %A Duo Cheng %A Xingyu Zhou %A Bo Ji %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cheng26s %I PMLR %P 19179--19210 %U https://proceedings.mlr.press/v306/cheng26s.html %V 306 %X We study hybrid reinforcement learning (RL) in adversarial Markov Decision Processes (MDPs), where the learner simultaneously receives on-policy feedback from the executed policy and off-policy feedback from a fixed behavior policy, and loss functions can change arbitrarily over time. On-policy feedback allows exploration and ensures the worst-case guarantee against any comparator policy, while off-policy feedback provides coverage-dependent guarantee that scales with the "mismatch" between the behavior and comparator policies (called coverage ratio) and can be sharper than on-policy results whenever this ratio is small. We propose a new hybrid RL framework that accommodates adversarial losses and unknown transitions, preserving off-policy guarantees while ensuring non-trivial worst-case performance.
APA
Cheng, D., Zhou, X. & Ji, B.. (2026). Hybrid Reinforcement Learning in Adversarial Markov Decision Processes. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:19179-19210 Available from https://proceedings.mlr.press/v306/cheng26s.html.

Related Material