Model-Free Reinforcement Learning with Skew-Symmetric Bilinear Utilities

Hugo Gilbert LIP6-UPMC, Bruno Zanuttini, Paul Weng SYSU-CMU JIE, Paolo Viappiani Lip6 Paris, Esther Nicart Cordon Electronics DS2i
Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, PMLR R14:276-285, 2016.

Abstract

In reinforcement learning, policies are typically evaluated according to the expectation of cumulated rewards. Researchers in decision theory have argued that more sophisticated decision criteria can better model the preferences of a decision maker. In particular, Skew-Symmetric Bilinear (SSB) utility functions generalize vonNeumann and Morgenstern’s expected utility (EU) theory to encompass rational decision behaviors that EU cannot accommodate. In this paper, we adopt an SSB utility function to compare policies in the reinforcement learning setting. We provide a model-free SSB reinforcement learning algorithm, SSB Q-learning, and prove its convergence towards a policy that is epsilon-optimal according to SSB. The proposed algorithm is an adaptation of fictitious play [Brown, 1951] combined with techniques from stochastic approximation [Borkar, 1997]. We also present some experimental results which evaluate our approach in a variety of settings.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR14-lip6-upmc16a, title = {Model-Free Reinforcement Learning with Skew-Symmetric Bilinear Utilities}, author = {LIP6-UPMC, Hugo Gilbert and Zanuttini, Bruno and JIE, Paul Weng SYSU-CMU and Paris, Paolo Viappiani Lip6 and DS2i, Esther Nicart Cordon Electronics}, booktitle = {Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence}, pages = {276--285}, year = {2016}, editor = {Ihler, Alexander and Janzing, Dominik}, volume = {R14}, series = {Proceedings of Machine Learning Research}, month = {25--29 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r14/main/assets/lip6-upmc16a/lip6-upmc16a.pdf}, url = {https://proceedings.mlr.press/r14/lip6-upmc16a.html}, abstract = {In reinforcement learning, policies are typically evaluated according to the expectation of cumulated rewards. Researchers in decision theory have argued that more sophisticated decision criteria can better model the preferences of a decision maker. In particular, Skew-Symmetric Bilinear (SSB) utility functions generalize vonNeumann and Morgenstern’s expected utility (EU) theory to encompass rational decision behaviors that EU cannot accommodate. In this paper, we adopt an SSB utility function to compare policies in the reinforcement learning setting. We provide a model-free SSB reinforcement learning algorithm, SSB Q-learning, and prove its convergence towards a policy that is epsilon-optimal according to SSB. The proposed algorithm is an adaptation of fictitious play [Brown, 1951] combined with techniques from stochastic approximation [Borkar, 1997]. We also present some experimental results which evaluate our approach in a variety of settings.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Model-Free Reinforcement Learning with Skew-Symmetric Bilinear Utilities %A Hugo Gilbert LIP6-UPMC %A Bruno Zanuttini %A Paul Weng SYSU-CMU JIE %A Paolo Viappiani Lip6 Paris %A Esther Nicart Cordon Electronics DS2i %B Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2016 %E Alexander Ihler %E Dominik Janzing %F pmlr-vR14-lip6-upmc16a %I PMLR %P 276--285 %U https://proceedings.mlr.press/r14/lip6-upmc16a.html %V R14 %X In reinforcement learning, policies are typically evaluated according to the expectation of cumulated rewards. Researchers in decision theory have argued that more sophisticated decision criteria can better model the preferences of a decision maker. In particular, Skew-Symmetric Bilinear (SSB) utility functions generalize vonNeumann and Morgenstern’s expected utility (EU) theory to encompass rational decision behaviors that EU cannot accommodate. In this paper, we adopt an SSB utility function to compare policies in the reinforcement learning setting. We provide a model-free SSB reinforcement learning algorithm, SSB Q-learning, and prove its convergence towards a policy that is epsilon-optimal according to SSB. The proposed algorithm is an adaptation of fictitious play [Brown, 1951] combined with techniques from stochastic approximation [Borkar, 1997]. We also present some experimental results which evaluate our approach in a variety of settings. %Z Reissued by PMLR on 04 October 2026.
APA
LIP6-UPMC, H.G., Zanuttini, B., JIE, P.W.S., Paris, P.V.L. & DS2i, E.N.C.E.. (2016). Model-Free Reinforcement Learning with Skew-Symmetric Bilinear Utilities. Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R14:276-285 Available from https://proceedings.mlr.press/r14/lip6-upmc16a.html. Reissued by PMLR on 04 October 2026.

Related Material