Curriculum Reinforcement Learning for Black-Box Prompt Tuning via Large Language Models

Shuai Gong, Chaoran Cui, Xiaolin Dong, Chunyun Zhang, Linwei Fan
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:35925-35942, 2026.

Abstract

Black-box prompt tuning (BBPT) aims to optimize input prompts for large models where internal parameters and gradients are inaccessible. However, existing methods fail to simultaneously address the dual challenges of prompt interpretability and query efficiency. To address these challenges, we propose CRL-BPT, a curriculum reinforcement learning framework that utilizes a large language model as an agent to generate human-readable prompts. Specifically, CRL-BPT implements a dynamic curriculum schedule on two auxiliary objectives: an imitation loss and an innovation loss. By dynamically weighting these objectives, CRL-BPT regularizes the RL process, guiding the agent from mimicking reference prompts to discovering novel patterns. Additionally, we introduce tailored stabilization mechanisms comprising historical loss normalization and relative reward calibration to promote more stable training. Extensive experiments demonstrate that CRL-BPT establishes new state-of-the-art performance and generates highly interpretable prompts under a strict budget of API calls. Code is available at https://github.com/GongShuai8210/CRL-BPT.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gong26d, title = {Curriculum Reinforcement Learning for Black-Box Prompt Tuning via Large Language Models}, author = {Gong, Shuai and Cui, Chaoran and Dong, Xiaolin and Zhang, Chunyun and Fan, Linwei}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {35925--35942}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gong26d/gong26d.pdf}, url = {https://proceedings.mlr.press/v306/gong26d.html}, abstract = {Black-box prompt tuning (BBPT) aims to optimize input prompts for large models where internal parameters and gradients are inaccessible. However, existing methods fail to simultaneously address the dual challenges of prompt interpretability and query efficiency. To address these challenges, we propose CRL-BPT, a curriculum reinforcement learning framework that utilizes a large language model as an agent to generate human-readable prompts. Specifically, CRL-BPT implements a dynamic curriculum schedule on two auxiliary objectives: an imitation loss and an innovation loss. By dynamically weighting these objectives, CRL-BPT regularizes the RL process, guiding the agent from mimicking reference prompts to discovering novel patterns. Additionally, we introduce tailored stabilization mechanisms comprising historical loss normalization and relative reward calibration to promote more stable training. Extensive experiments demonstrate that CRL-BPT establishes new state-of-the-art performance and generates highly interpretable prompts under a strict budget of API calls. Code is available at https://github.com/GongShuai8210/CRL-BPT.} }
Endnote
%0 Conference Paper %T Curriculum Reinforcement Learning for Black-Box Prompt Tuning via Large Language Models %A Shuai Gong %A Chaoran Cui %A Xiaolin Dong %A Chunyun Zhang %A Linwei Fan %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gong26d %I PMLR %P 35925--35942 %U https://proceedings.mlr.press/v306/gong26d.html %V 306 %X Black-box prompt tuning (BBPT) aims to optimize input prompts for large models where internal parameters and gradients are inaccessible. However, existing methods fail to simultaneously address the dual challenges of prompt interpretability and query efficiency. To address these challenges, we propose CRL-BPT, a curriculum reinforcement learning framework that utilizes a large language model as an agent to generate human-readable prompts. Specifically, CRL-BPT implements a dynamic curriculum schedule on two auxiliary objectives: an imitation loss and an innovation loss. By dynamically weighting these objectives, CRL-BPT regularizes the RL process, guiding the agent from mimicking reference prompts to discovering novel patterns. Additionally, we introduce tailored stabilization mechanisms comprising historical loss normalization and relative reward calibration to promote more stable training. Extensive experiments demonstrate that CRL-BPT establishes new state-of-the-art performance and generates highly interpretable prompts under a strict budget of API calls. Code is available at https://github.com/GongShuai8210/CRL-BPT.
APA
Gong, S., Cui, C., Dong, X., Zhang, C. & Fan, L.. (2026). Curriculum Reinforcement Learning for Black-Box Prompt Tuning via Large Language Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:35925-35942 Available from https://proceedings.mlr.press/v306/gong26d.html.

Related Material