EchoRL: Reinforcement Learning via Rollout Echoing

Jinhe Bi,  Aniri, Minglai Yang, Xingcheng Zhou, Wenke Huang, Sikuan Yan, Yujun Wang, Zixuan Cao, Michael Färber, Xun Xiao, Volker Tresp, Yunpu Ma
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:7999-8021, 2026.

Abstract

Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models. However, as training proceeds, the learning signal can collapse thus makes the training gain become marginal and ineffective. Specifically, a growing fraction of prompts’ rollouts become advantage-degenerated: all the self-generated rollouts show verified-success, making the standard deviation over their rewards be zero; accordingly each rollout’s advantage becomes degenerated (zero) as well. Given such rollouts’ advantages, the policy-gradient for model optimization eventually vanishes, capping the training performance. We argue that some of these rollouts still contain valuable learning signals but unfortunately omitted with the existing RLVR methods. In this paper, inspired through analyzing the entropy pattern behind golden trajectories produced by external expert models, we propose EchoRL for better exploiting the advantage-degenerated rollouts to further improve the training performance. EchoRL is a lightweight module that first identifies an EchoClip from verified-success rollouts based on their step-level entropy values, and then feeds this clip back as an auxiliary supervision signal in the RL objective. Extensive experiments across 10 benchmarks, 5 LLM backbones, and 7 popular RLVR post-training methods demonstrate that EchoRL consistently improves RLVR post-training with minimal overhead.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bi26a, title = {{E}cho{RL}: Reinforcement Learning via Rollout Echoing}, author = {Bi, Jinhe and Aniri and Yang, Minglai and Zhou, Xingcheng and Huang, Wenke and Yan, Sikuan and Wang, Yujun and Cao, Zixuan and F\"{a}rber, Michael and Xiao, Xun and Tresp, Volker and Ma, Yunpu}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {7999--8021}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bi26a/bi26a.pdf}, url = {https://proceedings.mlr.press/v306/bi26a.html}, abstract = {Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models. However, as training proceeds, the learning signal can collapse thus makes the training gain become marginal and ineffective. Specifically, a growing fraction of prompts’ rollouts become advantage-degenerated: all the self-generated rollouts show verified-success, making the standard deviation over their rewards be zero; accordingly each rollout’s advantage becomes degenerated (zero) as well. Given such rollouts’ advantages, the policy-gradient for model optimization eventually vanishes, capping the training performance. We argue that some of these rollouts still contain valuable learning signals but unfortunately omitted with the existing RLVR methods. In this paper, inspired through analyzing the entropy pattern behind golden trajectories produced by external expert models, we propose EchoRL for better exploiting the advantage-degenerated rollouts to further improve the training performance. EchoRL is a lightweight module that first identifies an EchoClip from verified-success rollouts based on their step-level entropy values, and then feeds this clip back as an auxiliary supervision signal in the RL objective. Extensive experiments across 10 benchmarks, 5 LLM backbones, and 7 popular RLVR post-training methods demonstrate that EchoRL consistently improves RLVR post-training with minimal overhead.} }
Endnote
%0 Conference Paper %T EchoRL: Reinforcement Learning via Rollout Echoing %A Jinhe Bi %A Aniri %A Minglai Yang %A Xingcheng Zhou %A Wenke Huang %A Sikuan Yan %A Yujun Wang %A Zixuan Cao %A Michael Färber %A Xun Xiao %A Volker Tresp %A Yunpu Ma %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bi26a %I PMLR %P 7999--8021 %U https://proceedings.mlr.press/v306/bi26a.html %V 306 %X Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models. However, as training proceeds, the learning signal can collapse thus makes the training gain become marginal and ineffective. Specifically, a growing fraction of prompts’ rollouts become advantage-degenerated: all the self-generated rollouts show verified-success, making the standard deviation over their rewards be zero; accordingly each rollout’s advantage becomes degenerated (zero) as well. Given such rollouts’ advantages, the policy-gradient for model optimization eventually vanishes, capping the training performance. We argue that some of these rollouts still contain valuable learning signals but unfortunately omitted with the existing RLVR methods. In this paper, inspired through analyzing the entropy pattern behind golden trajectories produced by external expert models, we propose EchoRL for better exploiting the advantage-degenerated rollouts to further improve the training performance. EchoRL is a lightweight module that first identifies an EchoClip from verified-success rollouts based on their step-level entropy values, and then feeds this clip back as an auxiliary supervision signal in the RL objective. Extensive experiments across 10 benchmarks, 5 LLM backbones, and 7 popular RLVR post-training methods demonstrate that EchoRL consistently improves RLVR post-training with minimal overhead.
APA
Bi, J., Aniri, , Yang, M., Zhou, X., Huang, W., Yan, S., Wang, Y., Cao, Z., Färber, M., Xiao, X., Tresp, V. & Ma, Y.. (2026). EchoRL: Reinforcement Learning via Rollout Echoing. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:7999-8021 Available from https://proceedings.mlr.press/v306/bi26a.html.

Related Material