Evolution Strategies at the Hyperscale

Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Luca Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, Jakob Nicolaus Foerster
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:107670-107741, 2026.

Abstract

Evolution Strategies (ES) is a class of powerful black-box optimisation methods that are highly parallelisable and can handle non-differentiable and noisy objectives. However, naïve ES becomes prohibitively expensive at scale on GPUs due to the low arithmetic intensity of batched matrix multiplications with unstructured random perturbations. We introduce Evolution Guided GeneRal Optimisation via Low-rank Learning (EGGROLL), which improves arithmetic intensity by structuring individual perturbations as rank-$r$ matrices, resulting in a hundredfold increase in training speed for billion-parameter models at large population sizes, achieving up to 91% of the throughput of pure batch inference. We provide a rigorous theoretical analysis of ES for high-dimensional parameter objectives, investigating conditions needed for ES updates to converge in high dimensions, revealing a linearising effect, and proving consistency between EGGROLL and ES as parameter dimension increases. Our experiments show that EGGROLL: (1) enables the stable pretraining of nonlinear recurrent language models that operate purely in integer datatypes, (2) is competitive with GRPO for post-training LLMs on reasoning tasks, and (3) does not compromise performance compared to ES in tabula rasa RL settings, despite being faster.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-sarkar26a, title = {Evolution Strategies at the Hyperscale}, author = {Sarkar, Bidipta and Fellows, Mattie and Duque, Juan Agustin and Letcher, Alistair and Villares, Antonio Le\'{o}n and Sims, Anya and Wibault, Clarisse and Samsonov, Dmitry and Cope, Dylan and Liesen, Jarek Luca and Li, Kang and Seier, Lukas and Wolf, Theo and Berdica, Uljad and Mohl, Valentin and Goldie, Alexander David and Courville, Aaron and Sevegnani, Karin and Whiteson, Shimon and Foerster, Jakob Nicolaus}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {107670--107741}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/sarkar26a/sarkar26a.pdf}, url = {https://proceedings.mlr.press/v306/sarkar26a.html}, abstract = {Evolution Strategies (ES) is a class of powerful black-box optimisation methods that are highly parallelisable and can handle non-differentiable and noisy objectives. However, naïve ES becomes prohibitively expensive at scale on GPUs due to the low arithmetic intensity of batched matrix multiplications with unstructured random perturbations. We introduce Evolution Guided GeneRal Optimisation via Low-rank Learning (EGGROLL), which improves arithmetic intensity by structuring individual perturbations as rank-$r$ matrices, resulting in a hundredfold increase in training speed for billion-parameter models at large population sizes, achieving up to 91% of the throughput of pure batch inference. We provide a rigorous theoretical analysis of ES for high-dimensional parameter objectives, investigating conditions needed for ES updates to converge in high dimensions, revealing a linearising effect, and proving consistency between EGGROLL and ES as parameter dimension increases. Our experiments show that EGGROLL: (1) enables the stable pretraining of nonlinear recurrent language models that operate purely in integer datatypes, (2) is competitive with GRPO for post-training LLMs on reasoning tasks, and (3) does not compromise performance compared to ES in tabula rasa RL settings, despite being faster.} }
Endnote
%0 Conference Paper %T Evolution Strategies at the Hyperscale %A Bidipta Sarkar %A Mattie Fellows %A Juan Agustin Duque %A Alistair Letcher %A Antonio León Villares %A Anya Sims %A Clarisse Wibault %A Dmitry Samsonov %A Dylan Cope %A Jarek Luca Liesen %A Kang Li %A Lukas Seier %A Theo Wolf %A Uljad Berdica %A Valentin Mohl %A Alexander David Goldie %A Aaron Courville %A Karin Sevegnani %A Shimon Whiteson %A Jakob Nicolaus Foerster %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-sarkar26a %I PMLR %P 107670--107741 %U https://proceedings.mlr.press/v306/sarkar26a.html %V 306 %X Evolution Strategies (ES) is a class of powerful black-box optimisation methods that are highly parallelisable and can handle non-differentiable and noisy objectives. However, naïve ES becomes prohibitively expensive at scale on GPUs due to the low arithmetic intensity of batched matrix multiplications with unstructured random perturbations. We introduce Evolution Guided GeneRal Optimisation via Low-rank Learning (EGGROLL), which improves arithmetic intensity by structuring individual perturbations as rank-$r$ matrices, resulting in a hundredfold increase in training speed for billion-parameter models at large population sizes, achieving up to 91% of the throughput of pure batch inference. We provide a rigorous theoretical analysis of ES for high-dimensional parameter objectives, investigating conditions needed for ES updates to converge in high dimensions, revealing a linearising effect, and proving consistency between EGGROLL and ES as parameter dimension increases. Our experiments show that EGGROLL: (1) enables the stable pretraining of nonlinear recurrent language models that operate purely in integer datatypes, (2) is competitive with GRPO for post-training LLMs on reasoning tasks, and (3) does not compromise performance compared to ES in tabula rasa RL settings, despite being faster.
APA
Sarkar, B., Fellows, M., Duque, J.A., Letcher, A., Villares, A.L., Sims, A., Wibault, C., Samsonov, D., Cope, D., Liesen, J.L., Li, K., Seier, L., Wolf, T., Berdica, U., Mohl, V., Goldie, A.D., Courville, A., Sevegnani, K., Whiteson, S. & Foerster, J.N.. (2026). Evolution Strategies at the Hyperscale. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:107670-107741 Available from https://proceedings.mlr.press/v306/sarkar26a.html.

Related Material