General Bayesian Policy Learning

Masahiro Kato
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2792-2827, 2026.

Abstract

This study proposes the General {Bayes} framework for policy learning. We consider decision problems in which a decision-maker chooses an action from an action set to maximize its expected welfare. Typical examples include treatment choice and portfolio selection. In such problems, the statistical target is a decision rule, and the prediction of each outcome $Y(a)$ is not necessarily of primary interest. We formulate this policy learning problem by loss-based {Bayesian} updating. Our main technical device is a squared-loss surrogate for welfare maximization. We show that maximizing empirical welfare over a policy class is equivalent to minimizing a scaled squared error in the outcome difference, up to a quadratic regularization controlled by a tuning parameter $\zeta>0$. This rewriting yields a General {Bayes} posterior over decision rules that admits a {Gaussian} pseudo-likelihood interpretation. We clarify two {Bayesian} interpretations of the resulting generalized posterior, a working {Gaussian} view and a decision-theoretic loss-based view. As one implementation example, we introduce neural networks with tanh-squashed outputs. Finally, we provide theoretical guarantees in a {PAC-Bayes} style.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kato26a, title = {General {Bayesian} Policy Learning}, author = {Kato, Masahiro}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2792--2827}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kato26a/kato26a.pdf}, url = {https://proceedings.mlr.press/v337/kato26a.html}, abstract = {This study proposes the General {Bayes} framework for policy learning. We consider decision problems in which a decision-maker chooses an action from an action set to maximize its expected welfare. Typical examples include treatment choice and portfolio selection. In such problems, the statistical target is a decision rule, and the prediction of each outcome $Y(a)$ is not necessarily of primary interest. We formulate this policy learning problem by loss-based {Bayesian} updating. Our main technical device is a squared-loss surrogate for welfare maximization. We show that maximizing empirical welfare over a policy class is equivalent to minimizing a scaled squared error in the outcome difference, up to a quadratic regularization controlled by a tuning parameter $\zeta>0$. This rewriting yields a General {Bayes} posterior over decision rules that admits a {Gaussian} pseudo-likelihood interpretation. We clarify two {Bayesian} interpretations of the resulting generalized posterior, a working {Gaussian} view and a decision-theoretic loss-based view. As one implementation example, we introduce neural networks with tanh-squashed outputs. Finally, we provide theoretical guarantees in a {PAC-Bayes} style.} }
Endnote
%0 Conference Paper %T General Bayesian Policy Learning %A Masahiro Kato %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kato26a %I PMLR %P 2792--2827 %U https://proceedings.mlr.press/v337/kato26a.html %V 337 %X This study proposes the General {Bayes} framework for policy learning. We consider decision problems in which a decision-maker chooses an action from an action set to maximize its expected welfare. Typical examples include treatment choice and portfolio selection. In such problems, the statistical target is a decision rule, and the prediction of each outcome $Y(a)$ is not necessarily of primary interest. We formulate this policy learning problem by loss-based {Bayesian} updating. Our main technical device is a squared-loss surrogate for welfare maximization. We show that maximizing empirical welfare over a policy class is equivalent to minimizing a scaled squared error in the outcome difference, up to a quadratic regularization controlled by a tuning parameter $\zeta>0$. This rewriting yields a General {Bayes} posterior over decision rules that admits a {Gaussian} pseudo-likelihood interpretation. We clarify two {Bayesian} interpretations of the resulting generalized posterior, a working {Gaussian} view and a decision-theoretic loss-based view. As one implementation example, we introduce neural networks with tanh-squashed outputs. Finally, we provide theoretical guarantees in a {PAC-Bayes} style.
APA
Kato, M.. (2026). General Bayesian Policy Learning. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2792-2827 Available from https://proceedings.mlr.press/v337/kato26a.html.

Related Material