Towards Blackwell Optimality: Bellman Optimality Is All You Can Get

Victor Boone, Adrienne Tuynman
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1819-1827, 2026.

Abstract

Although average gain optimality is a commonly adopted performance measure in Markov Decision Processes (MDPs), it is often too asymptotic. Further incorporating measures of immediate losses leads to the hierarchy of bias optimalities, all the way up to Blackwell optimality. In this paper, we investigate the problem of identifying policies of such optimality orders. To that end, for each order, we construct a learning algorithm with vanishing probability of error. Furthermore, we characterize the class of MDPs for which identification algorithms can stop in finite time. That class corresponds to the MDPs with a unique Bellman optimal policy, and does not depend on the optimality order considered. Lastly, we provide a tractable stopping rule that when coupled to our learning algorithm triggers in finite time whenever it is possible to do so.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-boone26a, title = { Towards Blackwell Optimality: Bellman Optimality Is All You Can Get }, author = {Boone, Victor and Tuynman, Adrienne}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1819--1827}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/boone26a/boone26a.pdf}, url = {https://proceedings.mlr.press/v300/boone26a.html}, abstract = { Although average gain optimality is a commonly adopted performance measure in Markov Decision Processes (MDPs), it is often too asymptotic. Further incorporating measures of immediate losses leads to the hierarchy of bias optimalities, all the way up to Blackwell optimality. In this paper, we investigate the problem of identifying policies of such optimality orders. To that end, for each order, we construct a learning algorithm with vanishing probability of error. Furthermore, we characterize the class of MDPs for which identification algorithms can stop in finite time. That class corresponds to the MDPs with a unique Bellman optimal policy, and does not depend on the optimality order considered. Lastly, we provide a tractable stopping rule that when coupled to our learning algorithm triggers in finite time whenever it is possible to do so. } }
Endnote
%0 Conference Paper %T Towards Blackwell Optimality: Bellman Optimality Is All You Can Get %A Victor Boone %A Adrienne Tuynman %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-boone26a %I PMLR %P 1819--1827 %U https://proceedings.mlr.press/v300/boone26a.html %V 300 %X Although average gain optimality is a commonly adopted performance measure in Markov Decision Processes (MDPs), it is often too asymptotic. Further incorporating measures of immediate losses leads to the hierarchy of bias optimalities, all the way up to Blackwell optimality. In this paper, we investigate the problem of identifying policies of such optimality orders. To that end, for each order, we construct a learning algorithm with vanishing probability of error. Furthermore, we characterize the class of MDPs for which identification algorithms can stop in finite time. That class corresponds to the MDPs with a unique Bellman optimal policy, and does not depend on the optimality order considered. Lastly, we provide a tractable stopping rule that when coupled to our learning algorithm triggers in finite time whenever it is possible to do so.
APA
Boone, V. & Tuynman, A.. (2026). Towards Blackwell Optimality: Bellman Optimality Is All You Can Get . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1819-1827 Available from https://proceedings.mlr.press/v300/boone26a.html.

Related Material