Predicting evolutionary rate as a pretraining task improves genome language model representations

Micaela Elisa Consens, Kevin K Yang, James Brian Hall, Ashley Mae Conard, Bo Wang, Lorin Crawford, Alan M Moses, Alex Xijie Lu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:21349-21373, 2026.

Abstract

Genome language models (gLMs) have the potential to further understanding of regulatory genomics without requiring labeled data. Most gLMs are pretrained using sequence reconstruction tasks inspired by natural language processing, but recent studies have shown that these gLMs often fail to capture biological signal. To overcome this, we introduce pretraining tasks that predict the rate of evolution. These tasks are designed so that they can be composed with sequence reconstruction, enabling a controlled comparison of predicting sequence only, evolutionary rate only, or both. To address gaps in existing evaluations, we developed a suite of biologically grounded benchmarks. Across these tasks, and for established variant effect prediction benchmarks, models pretrained on both sequence and evolutionary rate outperform those trained on sequence alone, and training on evolutionary rate can make even the relatively small models in our work competitive with much larger existing gLMs for some tasks on the human genome. These results establish evolution as a key training target for genome-scale models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-consens26a, title = {Predicting evolutionary rate as a pretraining task improves genome language model representations}, author = {Consens, Micaela Elisa and Yang, Kevin K and Hall, James Brian and Conard, Ashley Mae and Wang, Bo and Crawford, Lorin and Moses, Alan M and Lu, Alex Xijie}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {21349--21373}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/consens26a/consens26a.pdf}, url = {https://proceedings.mlr.press/v306/consens26a.html}, abstract = {Genome language models (gLMs) have the potential to further understanding of regulatory genomics without requiring labeled data. Most gLMs are pretrained using sequence reconstruction tasks inspired by natural language processing, but recent studies have shown that these gLMs often fail to capture biological signal. To overcome this, we introduce pretraining tasks that predict the rate of evolution. These tasks are designed so that they can be composed with sequence reconstruction, enabling a controlled comparison of predicting sequence only, evolutionary rate only, or both. To address gaps in existing evaluations, we developed a suite of biologically grounded benchmarks. Across these tasks, and for established variant effect prediction benchmarks, models pretrained on both sequence and evolutionary rate outperform those trained on sequence alone, and training on evolutionary rate can make even the relatively small models in our work competitive with much larger existing gLMs for some tasks on the human genome. These results establish evolution as a key training target for genome-scale models.} }
Endnote
%0 Conference Paper %T Predicting evolutionary rate as a pretraining task improves genome language model representations %A Micaela Elisa Consens %A Kevin K Yang %A James Brian Hall %A Ashley Mae Conard %A Bo Wang %A Lorin Crawford %A Alan M Moses %A Alex Xijie Lu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-consens26a %I PMLR %P 21349--21373 %U https://proceedings.mlr.press/v306/consens26a.html %V 306 %X Genome language models (gLMs) have the potential to further understanding of regulatory genomics without requiring labeled data. Most gLMs are pretrained using sequence reconstruction tasks inspired by natural language processing, but recent studies have shown that these gLMs often fail to capture biological signal. To overcome this, we introduce pretraining tasks that predict the rate of evolution. These tasks are designed so that they can be composed with sequence reconstruction, enabling a controlled comparison of predicting sequence only, evolutionary rate only, or both. To address gaps in existing evaluations, we developed a suite of biologically grounded benchmarks. Across these tasks, and for established variant effect prediction benchmarks, models pretrained on both sequence and evolutionary rate outperform those trained on sequence alone, and training on evolutionary rate can make even the relatively small models in our work competitive with much larger existing gLMs for some tasks on the human genome. These results establish evolution as a key training target for genome-scale models.
APA
Consens, M.E., Yang, K.K., Hall, J.B., Conard, A.M., Wang, B., Crawford, L., Moses, A.M. & Lu, A.X.. (2026). Predicting evolutionary rate as a pretraining task improves genome language model representations. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:21349-21373 Available from https://proceedings.mlr.press/v306/consens26a.html.

Related Material