A Covering Framework for Offline POMDPs Learning Using Belief Space Metric

Youheng Zhu, Yiping Lu
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1423-1431, 2026.

Abstract

In off-policy evaluation (OPE) for partially observable Markov decision processes (POMDPs), an agent must infer hidden states from past observations, which exacerbates both the curse of horizon and the curse of memory in existing OPE methods. This paper introduces a novel covering analysis framework that exploits the intrinsic metric structure of the belief space (distributions over latent states) to relax traditional coverage assumptions. By focusing on the policies with stability property, we derive error bounds that mitigate exponential blow-ups in horizon and memory length. Our unified analysis technique applies to a broad class of OPE algorithms, yielding concrete error bounds and coverage requirements expressed in terms of belief space metrics rather than raw history coverage. We illustrate the improved sample efficiency of this framework via case studies: the double sampling Bellman error minimization algorithm, and the memory-based future-dependent value functions (FDVF). In both cases, our coverage definition based on the belief-space metric yields tighter bounds.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-zhu26a, title = { A Covering Framework for Offline POMDPs Learning Using Belief Space Metric }, author = {Zhu, Youheng and Lu, Yiping}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1423--1431}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/zhu26a/zhu26a.pdf}, url = {https://proceedings.mlr.press/v300/zhu26a.html}, abstract = { In off-policy evaluation (OPE) for partially observable Markov decision processes (POMDPs), an agent must infer hidden states from past observations, which exacerbates both the curse of horizon and the curse of memory in existing OPE methods. This paper introduces a novel covering analysis framework that exploits the intrinsic metric structure of the belief space (distributions over latent states) to relax traditional coverage assumptions. By focusing on the policies with stability property, we derive error bounds that mitigate exponential blow-ups in horizon and memory length. Our unified analysis technique applies to a broad class of OPE algorithms, yielding concrete error bounds and coverage requirements expressed in terms of belief space metrics rather than raw history coverage. We illustrate the improved sample efficiency of this framework via case studies: the double sampling Bellman error minimization algorithm, and the memory-based future-dependent value functions (FDVF). In both cases, our coverage definition based on the belief-space metric yields tighter bounds. } }
Endnote
%0 Conference Paper %T A Covering Framework for Offline POMDPs Learning Using Belief Space Metric %A Youheng Zhu %A Yiping Lu %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-zhu26a %I PMLR %P 1423--1431 %U https://proceedings.mlr.press/v300/zhu26a.html %V 300 %X In off-policy evaluation (OPE) for partially observable Markov decision processes (POMDPs), an agent must infer hidden states from past observations, which exacerbates both the curse of horizon and the curse of memory in existing OPE methods. This paper introduces a novel covering analysis framework that exploits the intrinsic metric structure of the belief space (distributions over latent states) to relax traditional coverage assumptions. By focusing on the policies with stability property, we derive error bounds that mitigate exponential blow-ups in horizon and memory length. Our unified analysis technique applies to a broad class of OPE algorithms, yielding concrete error bounds and coverage requirements expressed in terms of belief space metrics rather than raw history coverage. We illustrate the improved sample efficiency of this framework via case studies: the double sampling Bellman error minimization algorithm, and the memory-based future-dependent value functions (FDVF). In both cases, our coverage definition based on the belief-space metric yields tighter bounds.
APA
Zhu, Y. & Lu, Y.. (2026). A Covering Framework for Offline POMDPs Learning Using Belief Space Metric . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1423-1431 Available from https://proceedings.mlr.press/v300/zhu26a.html.

Related Material