Tournament Style RL: Stabilizing Policy Optimization on Non Verifiable Problems

Gurusha Juneja, Shubham Milind Phal, Jennifer She, Lisa Wang, Dorsa Sadigh, Anca Dragan, William Yang Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:54871-54887, 2026.

Abstract

Many real-world tasks are non-verifiable—there is no objective ground truth, and quality must be judged subjectively—making reward design for RL difficult. Existing approaches based on scalar rubric scores or single comparisons are often noisy, poorly calibrated, or provide sparse learning signals. We introduce Tournament Style RL (TSRL), which constructs rewards from rubric-guided pairwise judgments against a fixed set of anchor responses, using win-rate as the reward for policy optimization. This aggregation of comparisons against anchor responses yields a signal that is more robust to the judge noise by stabilizing the reference frame, reducing the variance in reward. We test across four non-verifiable tasks and two backbone LLMs, and find that TSRL improves average win-rate by $+43.8$ points over the base model and $+22.8$ points over the strongest baseline. TSRL scales with the number of anchors, remains robust under weak or partially corrupted judges, the results are supported by blinded human preference studies.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-juneja26b, title = {Tournament Style {RL}: Stabilizing Policy Optimization on Non Verifiable Problems}, author = {Juneja, Gurusha and Phal, Shubham Milind and She, Jennifer and Wang, Lisa and Sadigh, Dorsa and Dragan, Anca and Wang, William Yang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {54871--54887}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/juneja26b/juneja26b.pdf}, url = {https://proceedings.mlr.press/v306/juneja26b.html}, abstract = {Many real-world tasks are non-verifiable—there is no objective ground truth, and quality must be judged subjectively—making reward design for RL difficult. Existing approaches based on scalar rubric scores or single comparisons are often noisy, poorly calibrated, or provide sparse learning signals. We introduce Tournament Style RL (TSRL), which constructs rewards from rubric-guided pairwise judgments against a fixed set of anchor responses, using win-rate as the reward for policy optimization. This aggregation of comparisons against anchor responses yields a signal that is more robust to the judge noise by stabilizing the reference frame, reducing the variance in reward. We test across four non-verifiable tasks and two backbone LLMs, and find that TSRL improves average win-rate by $+43.8$ points over the base model and $+22.8$ points over the strongest baseline. TSRL scales with the number of anchors, remains robust under weak or partially corrupted judges, the results are supported by blinded human preference studies.} }
Endnote
%0 Conference Paper %T Tournament Style RL: Stabilizing Policy Optimization on Non Verifiable Problems %A Gurusha Juneja %A Shubham Milind Phal %A Jennifer She %A Lisa Wang %A Dorsa Sadigh %A Anca Dragan %A William Yang Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-juneja26b %I PMLR %P 54871--54887 %U https://proceedings.mlr.press/v306/juneja26b.html %V 306 %X Many real-world tasks are non-verifiable—there is no objective ground truth, and quality must be judged subjectively—making reward design for RL difficult. Existing approaches based on scalar rubric scores or single comparisons are often noisy, poorly calibrated, or provide sparse learning signals. We introduce Tournament Style RL (TSRL), which constructs rewards from rubric-guided pairwise judgments against a fixed set of anchor responses, using win-rate as the reward for policy optimization. This aggregation of comparisons against anchor responses yields a signal that is more robust to the judge noise by stabilizing the reference frame, reducing the variance in reward. We test across four non-verifiable tasks and two backbone LLMs, and find that TSRL improves average win-rate by $+43.8$ points over the base model and $+22.8$ points over the strongest baseline. TSRL scales with the number of anchors, remains robust under weak or partially corrupted judges, the results are supported by blinded human preference studies.
APA
Juneja, G., Phal, S.M., She, J., Wang, L., Sadigh, D., Dragan, A. & Wang, W.Y.. (2026). Tournament Style RL: Stabilizing Policy Optimization on Non Verifiable Problems. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:54871-54887 Available from https://proceedings.mlr.press/v306/juneja26b.html.

Related Material