Duality-based Residual Estimation for Fully Offline Value-based Reinforcement Learning

Kohei Miyaguchi
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1576-1584, 2026.

Abstract

Value-based reinforcement learning (RL) efficiently handles high-dimensional state spaces, but existing methods lack a principled method for hyperparameter tuning without online interaction, limiting use in safety-critical and data-scarce domains. We propose the \textbf{Duality-based Residual Estimator (DRE)}, a simple offline validation metric for value-based offline RL. DRE is compatible with standard value-based Off-Policy Evaluation (OPE) and enables automatic hyperparameter selection, which is formalized through an adaptive extension of the Probably Approximately Correct (PAC) guarantee for Q-function selection. Our results address a key theoretical bottleneck toward \emph{fully offline} value-based RL, which enables deployment without extensive online tuning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-miyaguchi26a, title = { Duality-based Residual Estimation for Fully Offline Value-based Reinforcement Learning }, author = {Miyaguchi, Kohei}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1576--1584}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/miyaguchi26a/miyaguchi26a.pdf}, url = {https://proceedings.mlr.press/v300/miyaguchi26a.html}, abstract = { Value-based reinforcement learning (RL) efficiently handles high-dimensional state spaces, but existing methods lack a principled method for hyperparameter tuning without online interaction, limiting use in safety-critical and data-scarce domains. We propose the \textbf{Duality-based Residual Estimator (DRE)}, a simple offline validation metric for value-based offline RL. DRE is compatible with standard value-based Off-Policy Evaluation (OPE) and enables automatic hyperparameter selection, which is formalized through an adaptive extension of the Probably Approximately Correct (PAC) guarantee for Q-function selection. Our results address a key theoretical bottleneck toward \emph{fully offline} value-based RL, which enables deployment without extensive online tuning. } }
Endnote
%0 Conference Paper %T Duality-based Residual Estimation for Fully Offline Value-based Reinforcement Learning %A Kohei Miyaguchi %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-miyaguchi26a %I PMLR %P 1576--1584 %U https://proceedings.mlr.press/v300/miyaguchi26a.html %V 300 %X Value-based reinforcement learning (RL) efficiently handles high-dimensional state spaces, but existing methods lack a principled method for hyperparameter tuning without online interaction, limiting use in safety-critical and data-scarce domains. We propose the \textbf{Duality-based Residual Estimator (DRE)}, a simple offline validation metric for value-based offline RL. DRE is compatible with standard value-based Off-Policy Evaluation (OPE) and enables automatic hyperparameter selection, which is formalized through an adaptive extension of the Probably Approximately Correct (PAC) guarantee for Q-function selection. Our results address a key theoretical bottleneck toward \emph{fully offline} value-based RL, which enables deployment without extensive online tuning.
APA
Miyaguchi, K.. (2026). Duality-based Residual Estimation for Fully Offline Value-based Reinforcement Learning . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1576-1584 Available from https://proceedings.mlr.press/v300/miyaguchi26a.html.

Related Material