TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating

Hang Ding, Dongqi Liu, Qiming Feng, Jian Li, Tong Lei, Jiafu Wu, Shuo Wang, Jiangning Zhang, Chengjie Wang, Yabiao Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:25018-25034, 2026.

Abstract

Reinforcement learning from verifiable rewards (RLVR) has become an important paradigm for enhancing the reasoning capabilities of large language models, while it also involves a persistent tradeoff between optimization stability and learning efficiency. Token-level importance weighting supports fine-grained credit assignment, but it often introduces high variance and unstable parameter updates, whereas sequence-level optimization provides more stable learning dynamics while failing to fully exploit informative local signals. We introduce Trust-Gated Policy Optimization (TGPO), an efficient policy optimization framework that integrates two complementary mechanisms, namely sequence anchors and information gates. TGPO aligns token-wise updates with a stable sequence-level reference, which reduces the influence of extreme local likelihood fluctuations on the gradient, and a trust-based information gate adaptively modulates the contribution of token-level signals. By retaining and reweighting gradients from imperfect trajectories rather than excluding them, TGPO improves gradient utilization and sample efficiency while maintaining stable optimization behavior. Empirical results across seven mathematical reasoning datasets and multiple model scales show that TGPO consistently enhances learning efficiency and overall performance in outcome-supervised reinforcement learning settings.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ding26i, title = {{TGPO}: Efficient Policy Optimization through Sequence Anchor and Information Gating}, author = {Ding, Hang and Liu, Dongqi and Feng, Qiming and Li, Jian and Lei, Tong and Wu, Jiafu and Wang, Shuo and Zhang, Jiangning and Wang, Chengjie and Wang, Yabiao}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {25018--25034}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ding26i/ding26i.pdf}, url = {https://proceedings.mlr.press/v306/ding26i.html}, abstract = {Reinforcement learning from verifiable rewards (RLVR) has become an important paradigm for enhancing the reasoning capabilities of large language models, while it also involves a persistent tradeoff between optimization stability and learning efficiency. Token-level importance weighting supports fine-grained credit assignment, but it often introduces high variance and unstable parameter updates, whereas sequence-level optimization provides more stable learning dynamics while failing to fully exploit informative local signals. We introduce Trust-Gated Policy Optimization (TGPO), an efficient policy optimization framework that integrates two complementary mechanisms, namely sequence anchors and information gates. TGPO aligns token-wise updates with a stable sequence-level reference, which reduces the influence of extreme local likelihood fluctuations on the gradient, and a trust-based information gate adaptively modulates the contribution of token-level signals. By retaining and reweighting gradients from imperfect trajectories rather than excluding them, TGPO improves gradient utilization and sample efficiency while maintaining stable optimization behavior. Empirical results across seven mathematical reasoning datasets and multiple model scales show that TGPO consistently enhances learning efficiency and overall performance in outcome-supervised reinforcement learning settings.} }
Endnote
%0 Conference Paper %T TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating %A Hang Ding %A Dongqi Liu %A Qiming Feng %A Jian Li %A Tong Lei %A Jiafu Wu %A Shuo Wang %A Jiangning Zhang %A Chengjie Wang %A Yabiao Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ding26i %I PMLR %P 25018--25034 %U https://proceedings.mlr.press/v306/ding26i.html %V 306 %X Reinforcement learning from verifiable rewards (RLVR) has become an important paradigm for enhancing the reasoning capabilities of large language models, while it also involves a persistent tradeoff between optimization stability and learning efficiency. Token-level importance weighting supports fine-grained credit assignment, but it often introduces high variance and unstable parameter updates, whereas sequence-level optimization provides more stable learning dynamics while failing to fully exploit informative local signals. We introduce Trust-Gated Policy Optimization (TGPO), an efficient policy optimization framework that integrates two complementary mechanisms, namely sequence anchors and information gates. TGPO aligns token-wise updates with a stable sequence-level reference, which reduces the influence of extreme local likelihood fluctuations on the gradient, and a trust-based information gate adaptively modulates the contribution of token-level signals. By retaining and reweighting gradients from imperfect trajectories rather than excluding them, TGPO improves gradient utilization and sample efficiency while maintaining stable optimization behavior. Empirical results across seven mathematical reasoning datasets and multiple model scales show that TGPO consistently enhances learning efficiency and overall performance in outcome-supervised reinforcement learning settings.
APA
Ding, H., Liu, D., Feng, Q., Li, J., Lei, T., Wu, J., Wang, S., Zhang, J., Wang, C. & Wang, Y.. (2026). TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:25018-25034 Available from https://proceedings.mlr.press/v306/ding26i.html.

Related Material