GPT-Driven Drug Optimization with Structured Policy Optimization Post-training

Xuefeng Liu, Songhao Jiang, Siyu Chen, Zhuoran Yang, Yuxin Chen, Ian T. Foster, Rick L. Stevens
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1208-1242, 2026.

Abstract

Post-training is essential for steering generative models toward specific objectives. However, despite the growing importance of drug optimization, reinforcement learning algorithms tailored to this setting remain underexplored. In this work, we study the problem of drug optimization and propose a novel reinforcement learning framework for post-training a drug-optimization Generative Pre-trained Transformer (GPT). Our approach improves candidate molecules with respect to target objectives while preserving the desirable chemical properties of the original compounds. This work consists of two main components. (1) DrugImproverGPT, a framework designed to enhance the robustness and efficiency of drug optimization. It couples a GPT-based generative model with a theoretically grounded Structured Policy Optimization (SPO) algorithm. SPO provides a principled perspective on post-training generative models by explicitly aligning improvements in generated molecules with their corresponding input molecules under specified objectives. (2) A large-scale dataset comprising one million compounds, each annotated with OEDOCK docking scores across five human cancer-related proteins and 24 binding sites from the SARS-CoV-2 virus. Extensive in-silico experiments demonstrate that SPO generates candidates with improved values under the specified computational objectives. DrugImproverGPT is intended as a computational hit-to-lead prioritization framework whose generated candidates require subsequent experimental validation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-liu26c, title = {GPT-Driven Drug Optimization with Structured Policy Optimization Post-training}, author = {Liu, Xuefeng and Jiang, Songhao and Chen, Siyu and Yang, Zhuoran and Chen, Yuxin and Foster, Ian T. and Stevens, Rick L.}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1208--1242}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/liu26c/liu26c.pdf}, url = {https://proceedings.mlr.press/v340/liu26c.html}, abstract = {Post-training is essential for steering generative models toward specific objectives. However, despite the growing importance of drug optimization, reinforcement learning algorithms tailored to this setting remain underexplored. In this work, we study the problem of drug optimization and propose a novel reinforcement learning framework for post-training a drug-optimization Generative Pre-trained Transformer (GPT). Our approach improves candidate molecules with respect to target objectives while preserving the desirable chemical properties of the original compounds. This work consists of two main components. (1) DrugImproverGPT, a framework designed to enhance the robustness and efficiency of drug optimization. It couples a GPT-based generative model with a theoretically grounded Structured Policy Optimization (SPO) algorithm. SPO provides a principled perspective on post-training generative models by explicitly aligning improvements in generated molecules with their corresponding input molecules under specified objectives. (2) A large-scale dataset comprising one million compounds, each annotated with OEDOCK docking scores across five human cancer-related proteins and 24 binding sites from the SARS-CoV-2 virus. Extensive in-silico experiments demonstrate that SPO generates candidates with improved values under the specified computational objectives. DrugImproverGPT is intended as a computational hit-to-lead prioritization framework whose generated candidates require subsequent experimental validation.} }
Endnote
%0 Conference Paper %T GPT-Driven Drug Optimization with Structured Policy Optimization Post-training %A Xuefeng Liu %A Songhao Jiang %A Siyu Chen %A Zhuoran Yang %A Yuxin Chen %A Ian T. Foster %A Rick L. Stevens %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-liu26c %I PMLR %P 1208--1242 %U https://proceedings.mlr.press/v340/liu26c.html %V 340 %X Post-training is essential for steering generative models toward specific objectives. However, despite the growing importance of drug optimization, reinforcement learning algorithms tailored to this setting remain underexplored. In this work, we study the problem of drug optimization and propose a novel reinforcement learning framework for post-training a drug-optimization Generative Pre-trained Transformer (GPT). Our approach improves candidate molecules with respect to target objectives while preserving the desirable chemical properties of the original compounds. This work consists of two main components. (1) DrugImproverGPT, a framework designed to enhance the robustness and efficiency of drug optimization. It couples a GPT-based generative model with a theoretically grounded Structured Policy Optimization (SPO) algorithm. SPO provides a principled perspective on post-training generative models by explicitly aligning improvements in generated molecules with their corresponding input molecules under specified objectives. (2) A large-scale dataset comprising one million compounds, each annotated with OEDOCK docking scores across five human cancer-related proteins and 24 binding sites from the SARS-CoV-2 virus. Extensive in-silico experiments demonstrate that SPO generates candidates with improved values under the specified computational objectives. DrugImproverGPT is intended as a computational hit-to-lead prioritization framework whose generated candidates require subsequent experimental validation.
APA
Liu, X., Jiang, S., Chen, S., Yang, Z., Chen, Y., Foster, I.T. & Stevens, R.L.. (2026). GPT-Driven Drug Optimization with Structured Policy Optimization Post-training. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1208-1242 Available from https://proceedings.mlr.press/v340/liu26c.html.

Related Material