PRM-PBE: Process Reward Model for Reinforcement Learning in Programming-by-Example

Yue Fang, Zhi Jin, Jie An, Hongshen Chen, Jiangmeng Li, Xiaohong Chen, Naijun Zhan
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:29188-29200, 2026.

Abstract

Programming-by-Example (PBE), as a typical few-shot inductive reasoning paradigm, aims to synthesize corresponding algorithms from a set of input-output examples. Although Large Language Models (LLMs) have demonstrated strong program synthesis potential, they still remain ineffective when handling complex PBE tasks. Specifically, LLMs often struggle to accurately grasp the underlying intent of examples, resulting in synthesized programs that either partially satisfy the examples or completely deviate from the target. To address these limitations, we introduce a process-supervised reinforcement learning method that provides fine-grained feedback during the synthesis process, improving the ability of LLMs to capture the intended behavior of provided examples. Firstly, we develop a reasoning tree construction method that is used to build a PBE process supervision dataset. Subsequently, we train a process reward model through preference learning to evaluate the effectiveness of reasoning steps. Finally, we introduce a curriculum learning strategy based on the difficulty of PBE tasks, using Proximal Policy Optimization (PPO) to optimize the model. Experimental results on representative PBE benchmarks show that our approach achieves an average pass rate of 56.61%, significantly outperforming the state-of-the-art baseline by 8.73%.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fang26j, title = {{PRM}-{PBE}: Process Reward Model for Reinforcement Learning in Programming-by-Example}, author = {Fang, Yue and Jin, Zhi and An, Jie and Chen, Hongshen and Li, Jiangmeng and Chen, Xiaohong and Zhan, Naijun}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {29188--29200}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fang26j/fang26j.pdf}, url = {https://proceedings.mlr.press/v306/fang26j.html}, abstract = {Programming-by-Example (PBE), as a typical few-shot inductive reasoning paradigm, aims to synthesize corresponding algorithms from a set of input-output examples. Although Large Language Models (LLMs) have demonstrated strong program synthesis potential, they still remain ineffective when handling complex PBE tasks. Specifically, LLMs often struggle to accurately grasp the underlying intent of examples, resulting in synthesized programs that either partially satisfy the examples or completely deviate from the target. To address these limitations, we introduce a process-supervised reinforcement learning method that provides fine-grained feedback during the synthesis process, improving the ability of LLMs to capture the intended behavior of provided examples. Firstly, we develop a reasoning tree construction method that is used to build a PBE process supervision dataset. Subsequently, we train a process reward model through preference learning to evaluate the effectiveness of reasoning steps. Finally, we introduce a curriculum learning strategy based on the difficulty of PBE tasks, using Proximal Policy Optimization (PPO) to optimize the model. Experimental results on representative PBE benchmarks show that our approach achieves an average pass rate of 56.61%, significantly outperforming the state-of-the-art baseline by 8.73%.} }
Endnote
%0 Conference Paper %T PRM-PBE: Process Reward Model for Reinforcement Learning in Programming-by-Example %A Yue Fang %A Zhi Jin %A Jie An %A Hongshen Chen %A Jiangmeng Li %A Xiaohong Chen %A Naijun Zhan %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fang26j %I PMLR %P 29188--29200 %U https://proceedings.mlr.press/v306/fang26j.html %V 306 %X Programming-by-Example (PBE), as a typical few-shot inductive reasoning paradigm, aims to synthesize corresponding algorithms from a set of input-output examples. Although Large Language Models (LLMs) have demonstrated strong program synthesis potential, they still remain ineffective when handling complex PBE tasks. Specifically, LLMs often struggle to accurately grasp the underlying intent of examples, resulting in synthesized programs that either partially satisfy the examples or completely deviate from the target. To address these limitations, we introduce a process-supervised reinforcement learning method that provides fine-grained feedback during the synthesis process, improving the ability of LLMs to capture the intended behavior of provided examples. Firstly, we develop a reasoning tree construction method that is used to build a PBE process supervision dataset. Subsequently, we train a process reward model through preference learning to evaluate the effectiveness of reasoning steps. Finally, we introduce a curriculum learning strategy based on the difficulty of PBE tasks, using Proximal Policy Optimization (PPO) to optimize the model. Experimental results on representative PBE benchmarks show that our approach achieves an average pass rate of 56.61%, significantly outperforming the state-of-the-art baseline by 8.73%.
APA
Fang, Y., Jin, Z., An, J., Chen, H., Li, J., Chen, X. & Zhan, N.. (2026). PRM-PBE: Process Reward Model for Reinforcement Learning in Programming-by-Example. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:29188-29200 Available from https://proceedings.mlr.press/v306/fang26j.html.

Related Material