Best Policy Learning From Trajectory Preference Feedback

Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2116-2124, 2026.

Abstract

Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned reward models makes it vulnerable to mis-specification and reward hacking. Preference-based Reinforcement Learning (PbRL) offers a more robust alternative by directly leveraging noisy binary comparisons over trajectories. We study the best policy identification problem in PbRL, motivated by post-training optimization of generative models, for example, during multi-turn interactions. Learning in this setting combines an offline preference dataset—potentially biased or out-of-distribution and collected from a rater of subpar ’competence’—with online pure exploration, making systematic online learning essential. To this end, we propose Posterior Sampling for Preference Learning ($\mathsf{PSPL}$), a novel algorithm inspired by Top-Two Thompson Sampling that maintains posteriors over the reward model and dynamics. We provide the first Bayesian simple regret guarantees for PbRL and introduce an efficient approximation that outperforms existing baselines on simulation and image generation benchmarks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-agnihotri26a, title = { Best Policy Learning From Trajectory Preference Feedback }, author = {Agnihotri, Akhil and Jain, Rahul and Ramachandran, Deepak and Wen, Zheng}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2116--2124}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/agnihotri26a/agnihotri26a.pdf}, url = {https://proceedings.mlr.press/v300/agnihotri26a.html}, abstract = { Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned reward models makes it vulnerable to mis-specification and reward hacking. Preference-based Reinforcement Learning (PbRL) offers a more robust alternative by directly leveraging noisy binary comparisons over trajectories. We study the best policy identification problem in PbRL, motivated by post-training optimization of generative models, for example, during multi-turn interactions. Learning in this setting combines an offline preference dataset—potentially biased or out-of-distribution and collected from a rater of subpar ’competence’—with online pure exploration, making systematic online learning essential. To this end, we propose Posterior Sampling for Preference Learning ($\mathsf{PSPL}$), a novel algorithm inspired by Top-Two Thompson Sampling that maintains posteriors over the reward model and dynamics. We provide the first Bayesian simple regret guarantees for PbRL and introduce an efficient approximation that outperforms existing baselines on simulation and image generation benchmarks. } }
Endnote
%0 Conference Paper %T Best Policy Learning From Trajectory Preference Feedback %A Akhil Agnihotri %A Rahul Jain %A Deepak Ramachandran %A Zheng Wen %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-agnihotri26a %I PMLR %P 2116--2124 %U https://proceedings.mlr.press/v300/agnihotri26a.html %V 300 %X Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned reward models makes it vulnerable to mis-specification and reward hacking. Preference-based Reinforcement Learning (PbRL) offers a more robust alternative by directly leveraging noisy binary comparisons over trajectories. We study the best policy identification problem in PbRL, motivated by post-training optimization of generative models, for example, during multi-turn interactions. Learning in this setting combines an offline preference dataset—potentially biased or out-of-distribution and collected from a rater of subpar ’competence’—with online pure exploration, making systematic online learning essential. To this end, we propose Posterior Sampling for Preference Learning ($\mathsf{PSPL}$), a novel algorithm inspired by Top-Two Thompson Sampling that maintains posteriors over the reward model and dynamics. We provide the first Bayesian simple regret guarantees for PbRL and introduce an efficient approximation that outperforms existing baselines on simulation and image generation benchmarks.
APA
Agnihotri, A., Jain, R., Ramachandran, D. & Wen, Z.. (2026). Best Policy Learning From Trajectory Preference Feedback . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2116-2124 Available from https://proceedings.mlr.press/v300/agnihotri26a.html.

Related Material