Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM Reasoning

Zhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen, Weida Wang, Benteng Chen, Towsif Raiyan, Qingtao Pan, Yang Ouyang, Chaoda Song, Zhiqiang Gao, Shufei Zhang, Sumon Biswas
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:26269-26290, 2026.

Abstract

Large language models (LLMs) demonstrate strong reasoning abilities via Chain-of-Thought (CoT), but their token-level generation encourages local decisions and lacks global planning, often leading to redundant or inaccurate reasoning. Existing methods, such as tree-based search and reinforcement learning (RL), attempt to address this issue but incur high computational costs and still struggle to produce reliable reasoning trajectories. To address these challenges, we propose Plan-Then-Action Enhanced Reasoning with Group Relative Policy Optimization (PTA-GRPO), a two-stage framework designed to jointly improve high-level planning and fine-grained CoT reasoning. Specifically, in the first stage, a given LLM is responsible for summarizing CoT reasoning into compact high-level guidance, which is then leveraged for supervised fine-tuning. Then, we introduce a guidance-aware reinforcement learning method that jointly optimizes the final output and the quality of guidance, enhancing reasoning effectiveness. We evaluate PTA-GRPO on ten reasoning benchmarks across mathematics and natural sciences, using five diverse base models spanning multiple data modalities. The results show that PTA-GRPO consistently delivers stable and significant improvements across models and tasks, demonstrating strong effectiveness and generalization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-dou26c, title = {Plan Then Action: High-Level Planning Guidance Reinforcement Learning for {LLM} Reasoning}, author = {Dou, Zhihao and Zhao, Qinjian and Wan, Zhongwei and Dinggen, Zhang and Wang, Weida and Chen, Benteng and Raiyan, Towsif and Pan, Qingtao and Ouyang, Yang and Song, Chaoda and Gao, Zhiqiang and Zhang, Shufei and Biswas, Sumon}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {26269--26290}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/dou26c/dou26c.pdf}, url = {https://proceedings.mlr.press/v306/dou26c.html}, abstract = {Large language models (LLMs) demonstrate strong reasoning abilities via Chain-of-Thought (CoT), but their token-level generation encourages local decisions and lacks global planning, often leading to redundant or inaccurate reasoning. Existing methods, such as tree-based search and reinforcement learning (RL), attempt to address this issue but incur high computational costs and still struggle to produce reliable reasoning trajectories. To address these challenges, we propose Plan-Then-Action Enhanced Reasoning with Group Relative Policy Optimization (PTA-GRPO), a two-stage framework designed to jointly improve high-level planning and fine-grained CoT reasoning. Specifically, in the first stage, a given LLM is responsible for summarizing CoT reasoning into compact high-level guidance, which is then leveraged for supervised fine-tuning. Then, we introduce a guidance-aware reinforcement learning method that jointly optimizes the final output and the quality of guidance, enhancing reasoning effectiveness. We evaluate PTA-GRPO on ten reasoning benchmarks across mathematics and natural sciences, using five diverse base models spanning multiple data modalities. The results show that PTA-GRPO consistently delivers stable and significant improvements across models and tasks, demonstrating strong effectiveness and generalization.} }
Endnote
%0 Conference Paper %T Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM Reasoning %A Zhihao Dou %A Qinjian Zhao %A Zhongwei Wan %A Zhang Dinggen %A Weida Wang %A Benteng Chen %A Towsif Raiyan %A Qingtao Pan %A Yang Ouyang %A Chaoda Song %A Zhiqiang Gao %A Shufei Zhang %A Sumon Biswas %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-dou26c %I PMLR %P 26269--26290 %U https://proceedings.mlr.press/v306/dou26c.html %V 306 %X Large language models (LLMs) demonstrate strong reasoning abilities via Chain-of-Thought (CoT), but their token-level generation encourages local decisions and lacks global planning, often leading to redundant or inaccurate reasoning. Existing methods, such as tree-based search and reinforcement learning (RL), attempt to address this issue but incur high computational costs and still struggle to produce reliable reasoning trajectories. To address these challenges, we propose Plan-Then-Action Enhanced Reasoning with Group Relative Policy Optimization (PTA-GRPO), a two-stage framework designed to jointly improve high-level planning and fine-grained CoT reasoning. Specifically, in the first stage, a given LLM is responsible for summarizing CoT reasoning into compact high-level guidance, which is then leveraged for supervised fine-tuning. Then, we introduce a guidance-aware reinforcement learning method that jointly optimizes the final output and the quality of guidance, enhancing reasoning effectiveness. We evaluate PTA-GRPO on ten reasoning benchmarks across mathematics and natural sciences, using five diverse base models spanning multiple data modalities. The results show that PTA-GRPO consistently delivers stable and significant improvements across models and tasks, demonstrating strong effectiveness and generalization.
APA
Dou, Z., Zhao, Q., Wan, Z., Dinggen, Z., Wang, W., Chen, B., Raiyan, T., Pan, Q., Ouyang, Y., Song, C., Gao, Z., Zhang, S. & Biswas, S.. (2026). Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM Reasoning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:26269-26290 Available from https://proceedings.mlr.press/v306/dou26c.html.

Related Material