Escaping the Verifier: Learning to Reason via Demonstrations

Locke Cai, Max Ryabinin, Ivan Provilkov
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:10705-10735, 2026.

Abstract

Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) with task-specific verifiers. However, many real-world reasoning-intensive tasks lack verifiers, despite offering abundant expert demonstrations that remain under-utilized for reasoning-focused training. We introduce RARO (Relativistic Adversarial Reasoning Optimization), which learns strong reasoning capabilities from expert demonstrations alone via Inverse Reinforcement Learning. Our method sets up an adversarial game between a policy and a relativistic critic: the policy learns to mimic expert answers, while the critic aims to identify the expert among (expert, policy) answer pairs. Both the policy and the critic are trained jointly and continuously via RL, and we identify key stabilization techniques required for robust learning. Empirically, RARO significantly outperforms strong verifier-free baselines across all evaluation tasks: $+13.7%$ accuracy on Countdown ($1.5$B), +8.2% on DeepMath ($7$B), and $+19.1%$ win-rate on Poetry Writing ($7$B) against expert poems. RARO also exhibits similar robust scaling trends as RL with verifiers. These results demonstrate that RARO effectively elicits strong reasoning performance from expert demonstrations alone, enabling robust reasoning learning even when task-specific verifiers are unavailable.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cai26k, title = {Escaping the Verifier: Learning to Reason via Demonstrations}, author = {Cai, Locke and Ryabinin, Max and Provilkov, Ivan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {10705--10735}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cai26k/cai26k.pdf}, url = {https://proceedings.mlr.press/v306/cai26k.html}, abstract = {Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) with task-specific verifiers. However, many real-world reasoning-intensive tasks lack verifiers, despite offering abundant expert demonstrations that remain under-utilized for reasoning-focused training. We introduce RARO (Relativistic Adversarial Reasoning Optimization), which learns strong reasoning capabilities from expert demonstrations alone via Inverse Reinforcement Learning. Our method sets up an adversarial game between a policy and a relativistic critic: the policy learns to mimic expert answers, while the critic aims to identify the expert among (expert, policy) answer pairs. Both the policy and the critic are trained jointly and continuously via RL, and we identify key stabilization techniques required for robust learning. Empirically, RARO significantly outperforms strong verifier-free baselines across all evaluation tasks: $+13.7%$ accuracy on Countdown ($1.5$B), +8.2% on DeepMath ($7$B), and $+19.1%$ win-rate on Poetry Writing ($7$B) against expert poems. RARO also exhibits similar robust scaling trends as RL with verifiers. These results demonstrate that RARO effectively elicits strong reasoning performance from expert demonstrations alone, enabling robust reasoning learning even when task-specific verifiers are unavailable.} }
Endnote
%0 Conference Paper %T Escaping the Verifier: Learning to Reason via Demonstrations %A Locke Cai %A Max Ryabinin %A Ivan Provilkov %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cai26k %I PMLR %P 10705--10735 %U https://proceedings.mlr.press/v306/cai26k.html %V 306 %X Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) with task-specific verifiers. However, many real-world reasoning-intensive tasks lack verifiers, despite offering abundant expert demonstrations that remain under-utilized for reasoning-focused training. We introduce RARO (Relativistic Adversarial Reasoning Optimization), which learns strong reasoning capabilities from expert demonstrations alone via Inverse Reinforcement Learning. Our method sets up an adversarial game between a policy and a relativistic critic: the policy learns to mimic expert answers, while the critic aims to identify the expert among (expert, policy) answer pairs. Both the policy and the critic are trained jointly and continuously via RL, and we identify key stabilization techniques required for robust learning. Empirically, RARO significantly outperforms strong verifier-free baselines across all evaluation tasks: $+13.7%$ accuracy on Countdown ($1.5$B), +8.2% on DeepMath ($7$B), and $+19.1%$ win-rate on Poetry Writing ($7$B) against expert poems. RARO also exhibits similar robust scaling trends as RL with verifiers. These results demonstrate that RARO effectively elicits strong reasoning performance from expert demonstrations alone, enabling robust reasoning learning even when task-specific verifiers are unavailable.
APA
Cai, L., Ryabinin, M. & Provilkov, I.. (2026). Escaping the Verifier: Learning to Reason via Demonstrations. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:10705-10735 Available from https://proceedings.mlr.press/v306/cai26k.html.

Related Material