[edit]
GPT-Driven Drug Optimization with Structured Policy Optimization Post-training
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1208-1242, 2026.
Abstract
Post-training is essential for steering generative models toward specific objectives. However, despite the growing importance of drug optimization, reinforcement learning algorithms tailored to this setting remain underexplored. In this work, we study the problem of drug optimization and propose a novel reinforcement learning framework for post-training a drug-optimization Generative Pre-trained Transformer (GPT). Our approach improves candidate molecules with respect to target objectives while preserving the desirable chemical properties of the original compounds. This work consists of two main components. (1) DrugImproverGPT, a framework designed to enhance the robustness and efficiency of drug optimization. It couples a GPT-based generative model with a theoretically grounded Structured Policy Optimization (SPO) algorithm. SPO provides a principled perspective on post-training generative models by explicitly aligning improvements in generated molecules with their corresponding input molecules under specified objectives. (2) A large-scale dataset comprising one million compounds, each annotated with OEDOCK docking scores across five human cancer-related proteins and 24 binding sites from the SARS-CoV-2 virus. Extensive in-silico experiments demonstrate that SPO generates candidates with improved values under the specified computational objectives. DrugImproverGPT is intended as a computational hit-to-lead prioritization framework whose generated candidates require subsequent experimental validation.