RL-finetuning LLMs from on- and off-policy data with a single algorithm

Yunhao Tang, Taco Cohen, David W. Zhang, Gabriel Synnaeve, Rémi Munos
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:415-423, 2026.

Abstract

We introduce a novel reinforcement learning algorithm (AGRO, for Any-Generation Reward Optimization) for finetuning Large Language Models. AGRO leverages the concept of response consistency, which states that the optimal policy satisfies a notion of consistency across any possible generation of the model. We derive algorithms that find optimal solutions via sample-based policy gradient and provide theoretical guarantees on their convergence. Our experiments demonstrate the effectiveness of AGRO in both on-policy and off-policy settings, showing improved performance on the MATH dataset over baseline methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-tang26b, title = { RL-finetuning LLMs from on- and off-policy data with a single algorithm }, author = {Tang, Yunhao and Cohen, Taco and Zhang, David W. and Synnaeve, Gabriel and Munos, R{\'e}mi}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {415--423}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/tang26b/tang26b.pdf}, url = {https://proceedings.mlr.press/v300/tang26b.html}, abstract = { We introduce a novel reinforcement learning algorithm (AGRO, for Any-Generation Reward Optimization) for finetuning Large Language Models. AGRO leverages the concept of response consistency, which states that the optimal policy satisfies a notion of consistency across any possible generation of the model. We derive algorithms that find optimal solutions via sample-based policy gradient and provide theoretical guarantees on their convergence. Our experiments demonstrate the effectiveness of AGRO in both on-policy and off-policy settings, showing improved performance on the MATH dataset over baseline methods. } }
Endnote
%0 Conference Paper %T RL-finetuning LLMs from on- and off-policy data with a single algorithm %A Yunhao Tang %A Taco Cohen %A David W. Zhang %A Gabriel Synnaeve %A Rémi Munos %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-tang26b %I PMLR %P 415--423 %U https://proceedings.mlr.press/v300/tang26b.html %V 300 %X We introduce a novel reinforcement learning algorithm (AGRO, for Any-Generation Reward Optimization) for finetuning Large Language Models. AGRO leverages the concept of response consistency, which states that the optimal policy satisfies a notion of consistency across any possible generation of the model. We derive algorithms that find optimal solutions via sample-based policy gradient and provide theoretical guarantees on their convergence. Our experiments demonstrate the effectiveness of AGRO in both on-policy and off-policy settings, showing improved performance on the MATH dataset over baseline methods.
APA
Tang, Y., Cohen, T., Zhang, D.W., Synnaeve, G. & Munos, R.. (2026). RL-finetuning LLMs from on- and off-policy data with a single algorithm . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:415-423 Available from https://proceedings.mlr.press/v300/tang26b.html.

Related Material