Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Li Erran Li, Haokai Zhao, Jian Ma, Zijian Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu, Chiho Im, Seungju Han, Peng Xia, Tinson Xu, Yinxi Li, Deyao Zhu, Pheng-Ann Heng, Naoto Yokoya, Masashi Sugiyama, Jure Leskovec, Yejin Choi
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:137272-137305, 2026.

Abstract

Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and systematic reuse of biochemical knowledge. We introduce Proteo-R1, a reasoning-guided protein design framework that explicitly decouples molecular understanding from geometric generation. Proteo-R1 adopts a dual-expert architecture, where a multimodal large language model (LLM) serves as an understanding expert, analyzes protein sequences, structures, and textual context to identify key functional residues that govern binding and specificity. These residue-level decisions are then passed to a separate diffusion-based generation expert, which performs conditional co-design while respecting the fixed interaction anchors. This factorization mirrors how human experts approach molecular engineering: first, reasoning about critical interactions, then optimizing geometry subject to those constraints. By operationalizing reasoning as explicit residue-level commitments rather than latent textual guidance, Proteo-R1 achieves stable, interpretable, and modular integration of LLM reasoning with advanced geometric generative models. Code and demos are at https://proteor1.github.io.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-wu26bq, title = {Proteo-R1: Reasoning Foundation Models for De Novo Protein Design}, author = {Wu, Fang and Xuan, Weihao and Qi, Heli and Cao, Hanqun and Chang, Heng-Jui and Zhou, Zeqi and Li, Li Erran and Zhao, Haokai and Ma, Jian and Ma, Zijian Carl and Cheng, Yu-Chi and Pang, Kuan and Tang, Xiangru and Wang, Zehong and Li, Guanlue and Wang, Hanchen and Ying, Kejun and Lu, Pan and Im, Chiho and Han, Seungju and Xia, Peng and Xu, Tinson and Li, Yinxi and Zhu, Deyao and Heng, Pheng-Ann and Yokoya, Naoto and Sugiyama, Masashi and Leskovec, Jure and Choi, Yejin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {137272--137305}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/wu26bq/wu26bq.pdf}, url = {https://proceedings.mlr.press/v306/wu26bq.html}, abstract = {Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and systematic reuse of biochemical knowledge. We introduce Proteo-R1, a reasoning-guided protein design framework that explicitly decouples molecular understanding from geometric generation. Proteo-R1 adopts a dual-expert architecture, where a multimodal large language model (LLM) serves as an understanding expert, analyzes protein sequences, structures, and textual context to identify key functional residues that govern binding and specificity. These residue-level decisions are then passed to a separate diffusion-based generation expert, which performs conditional co-design while respecting the fixed interaction anchors. This factorization mirrors how human experts approach molecular engineering: first, reasoning about critical interactions, then optimizing geometry subject to those constraints. By operationalizing reasoning as explicit residue-level commitments rather than latent textual guidance, Proteo-R1 achieves stable, interpretable, and modular integration of LLM reasoning with advanced geometric generative models. Code and demos are at https://proteor1.github.io.} }
Endnote
%0 Conference Paper %T Proteo-R1: Reasoning Foundation Models for De Novo Protein Design %A Fang Wu %A Weihao Xuan %A Heli Qi %A Hanqun Cao %A Heng-Jui Chang %A Zeqi Zhou %A Li Erran Li %A Haokai Zhao %A Jian Ma %A Zijian Carl Ma %A Yu-Chi Cheng %A Kuan Pang %A Xiangru Tang %A Zehong Wang %A Guanlue Li %A Hanchen Wang %A Kejun Ying %A Pan Lu %A Chiho Im %A Seungju Han %A Peng Xia %A Tinson Xu %A Yinxi Li %A Deyao Zhu %A Pheng-Ann Heng %A Naoto Yokoya %A Masashi Sugiyama %A Jure Leskovec %A Yejin Choi %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-wu26bq %I PMLR %P 137272--137305 %U https://proceedings.mlr.press/v306/wu26bq.html %V 306 %X Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and systematic reuse of biochemical knowledge. We introduce Proteo-R1, a reasoning-guided protein design framework that explicitly decouples molecular understanding from geometric generation. Proteo-R1 adopts a dual-expert architecture, where a multimodal large language model (LLM) serves as an understanding expert, analyzes protein sequences, structures, and textual context to identify key functional residues that govern binding and specificity. These residue-level decisions are then passed to a separate diffusion-based generation expert, which performs conditional co-design while respecting the fixed interaction anchors. This factorization mirrors how human experts approach molecular engineering: first, reasoning about critical interactions, then optimizing geometry subject to those constraints. By operationalizing reasoning as explicit residue-level commitments rather than latent textual guidance, Proteo-R1 achieves stable, interpretable, and modular integration of LLM reasoning with advanced geometric generative models. Code and demos are at https://proteor1.github.io.
APA
Wu, F., Xuan, W., Qi, H., Cao, H., Chang, H., Zhou, Z., Li, L.E., Zhao, H., Ma, J., Ma, Z.C., Cheng, Y., Pang, K., Tang, X., Wang, Z., Li, G., Wang, H., Ying, K., Lu, P., Im, C., Han, S., Xia, P., Xu, T., Li, Y., Zhu, D., Heng, P., Yokoya, N., Sugiyama, M., Leskovec, J. & Choi, Y.. (2026). Proteo-R1: Reasoning Foundation Models for De Novo Protein Design. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:137272-137305 Available from https://proceedings.mlr.press/v306/wu26bq.html.

Related Material