FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications

Kieran Didi, Sarah Alamdari, Alex Xijie Lu, Bruce James Wittmann, Kadina E Johnston, Ava P Amini, Ali Madani, Maya Czeneszew, Christian Dallago, Kevin K Yang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:24724-24777, 2026.

Abstract

Machine learning methods that predict protein fitness from sequence remain sensitive to changes in data distributions, limiting generalization across common conditions encountered in protein engineering. Practically, protein engineers are thus left wondering about the effective utility of ML tools. The FLIP benchmark established protocols for testing generalization under some domain shifts, but it was limited to measurements of stability, binding, and viral capsid viability. We introduce FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns. Evaluating a suite of benchmark models across these datasets and suites reveals that simpler models often matched or outperformed fine-tuned protein language models on FLIP2, challenging the utility of existing transfer learning techniques. Provenance for all datasets has been recorded and we redistribute all data CC-BY 4.0 to facilitate continued progress.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-didi26a, title = {{FLIP}2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications}, author = {Didi, Kieran and Alamdari, Sarah and Lu, Alex Xijie and Wittmann, Bruce James and Johnston, Kadina E and Amini, Ava P and Madani, Ali and Czeneszew, Maya and Dallago, Christian and Yang, Kevin K}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {24724--24777}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/didi26a/didi26a.pdf}, url = {https://proceedings.mlr.press/v306/didi26a.html}, abstract = {Machine learning methods that predict protein fitness from sequence remain sensitive to changes in data distributions, limiting generalization across common conditions encountered in protein engineering. Practically, protein engineers are thus left wondering about the effective utility of ML tools. The FLIP benchmark established protocols for testing generalization under some domain shifts, but it was limited to measurements of stability, binding, and viral capsid viability. We introduce FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns. Evaluating a suite of benchmark models across these datasets and suites reveals that simpler models often matched or outperformed fine-tuned protein language models on FLIP2, challenging the utility of existing transfer learning techniques. Provenance for all datasets has been recorded and we redistribute all data CC-BY 4.0 to facilitate continued progress.} }
Endnote
%0 Conference Paper %T FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications %A Kieran Didi %A Sarah Alamdari %A Alex Xijie Lu %A Bruce James Wittmann %A Kadina E Johnston %A Ava P Amini %A Ali Madani %A Maya Czeneszew %A Christian Dallago %A Kevin K Yang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-didi26a %I PMLR %P 24724--24777 %U https://proceedings.mlr.press/v306/didi26a.html %V 306 %X Machine learning methods that predict protein fitness from sequence remain sensitive to changes in data distributions, limiting generalization across common conditions encountered in protein engineering. Practically, protein engineers are thus left wondering about the effective utility of ML tools. The FLIP benchmark established protocols for testing generalization under some domain shifts, but it was limited to measurements of stability, binding, and viral capsid viability. We introduce FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns. Evaluating a suite of benchmark models across these datasets and suites reveals that simpler models often matched or outperformed fine-tuned protein language models on FLIP2, challenging the utility of existing transfer learning techniques. Provenance for all datasets has been recorded and we redistribute all data CC-BY 4.0 to facilitate continued progress.
APA
Didi, K., Alamdari, S., Lu, A.X., Wittmann, B.J., Johnston, K.E., Amini, A.P., Madani, A., Czeneszew, M., Dallago, C. & Yang, K.K.. (2026). FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:24724-24777 Available from https://proceedings.mlr.press/v306/didi26a.html.

Related Material