Online Compatible Reward Identification from Preference Feedback

Simone Drago, Marco Mussi, Alberto Maria Metelli
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:26365-26389, 2026.

Abstract

In reinforcement learning, human preference feedback is emerging as a viable alternative to expert-designed reward functions, which can be difficult to engineer in real-world problems. However, despite the growing importance of preference feedback, how to effectively elicit preferences remains a fundamental open problem. This work focuses on the compatible reward identification task. The aim is to derive, starting from preference feedback, a reward function compatible with the observed preferences and accurate across the entire state-action space, ensuring higher transferability, safety, and interpretability. Indeed, the most common reinforcement learning from human feedback objective is to learn the optimal policy, requiring accuracy only in the portion of the state-action space that the agent visits. However, this goal cannot provide the same guarantees as compatible reward identification. First, we discuss commonalities and differences between the two goals. Then, we consider deterministic preferences, deriving the minimum number of interactions needed to identify the set of compatible rewards, and showing that using fewer queries may lead to arbitrarily large suboptimality. Finally, we focus on stochastic preferences generated via the Bradley-Terry (BT) model. We introduce the concepts of query basis and its index, relating them to the problem complexity. Upon this, we discuss the connection between the index of a basis and the BT model, as well as the limitations that the model induces in this setting. Additionally, we devise an algorithm to identify a nearly-optimal query basis with polynomial human query complexity.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-drago26a, title = {Online Compatible Reward Identification from Preference Feedback}, author = {Drago, Simone and Mussi, Marco and Metelli, Alberto Maria}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {26365--26389}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/drago26a/drago26a.pdf}, url = {https://proceedings.mlr.press/v306/drago26a.html}, abstract = {In reinforcement learning, human preference feedback is emerging as a viable alternative to expert-designed reward functions, which can be difficult to engineer in real-world problems. However, despite the growing importance of preference feedback, how to effectively elicit preferences remains a fundamental open problem. This work focuses on the compatible reward identification task. The aim is to derive, starting from preference feedback, a reward function compatible with the observed preferences and accurate across the entire state-action space, ensuring higher transferability, safety, and interpretability. Indeed, the most common reinforcement learning from human feedback objective is to learn the optimal policy, requiring accuracy only in the portion of the state-action space that the agent visits. However, this goal cannot provide the same guarantees as compatible reward identification. First, we discuss commonalities and differences between the two goals. Then, we consider deterministic preferences, deriving the minimum number of interactions needed to identify the set of compatible rewards, and showing that using fewer queries may lead to arbitrarily large suboptimality. Finally, we focus on stochastic preferences generated via the Bradley-Terry (BT) model. We introduce the concepts of query basis and its index, relating them to the problem complexity. Upon this, we discuss the connection between the index of a basis and the BT model, as well as the limitations that the model induces in this setting. Additionally, we devise an algorithm to identify a nearly-optimal query basis with polynomial human query complexity.} }
Endnote
%0 Conference Paper %T Online Compatible Reward Identification from Preference Feedback %A Simone Drago %A Marco Mussi %A Alberto Maria Metelli %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-drago26a %I PMLR %P 26365--26389 %U https://proceedings.mlr.press/v306/drago26a.html %V 306 %X In reinforcement learning, human preference feedback is emerging as a viable alternative to expert-designed reward functions, which can be difficult to engineer in real-world problems. However, despite the growing importance of preference feedback, how to effectively elicit preferences remains a fundamental open problem. This work focuses on the compatible reward identification task. The aim is to derive, starting from preference feedback, a reward function compatible with the observed preferences and accurate across the entire state-action space, ensuring higher transferability, safety, and interpretability. Indeed, the most common reinforcement learning from human feedback objective is to learn the optimal policy, requiring accuracy only in the portion of the state-action space that the agent visits. However, this goal cannot provide the same guarantees as compatible reward identification. First, we discuss commonalities and differences between the two goals. Then, we consider deterministic preferences, deriving the minimum number of interactions needed to identify the set of compatible rewards, and showing that using fewer queries may lead to arbitrarily large suboptimality. Finally, we focus on stochastic preferences generated via the Bradley-Terry (BT) model. We introduce the concepts of query basis and its index, relating them to the problem complexity. Upon this, we discuss the connection between the index of a basis and the BT model, as well as the limitations that the model induces in this setting. Additionally, we devise an algorithm to identify a nearly-optimal query basis with polynomial human query complexity.
APA
Drago, S., Mussi, M. & Metelli, A.M.. (2026). Online Compatible Reward Identification from Preference Feedback. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:26365-26389 Available from https://proceedings.mlr.press/v306/drago26a.html.

Related Material