Hierarchical Reinforcement Learning for Sparse-Reward Search in Commutative Algebra

Giorgi Butbaia, Paul Orland, Coco Huang, Davide Passaro, Lucas Fagan, Michele Tarquini, Hailong Dao, David Eisenbud, Ali Shehper, Sergei Gukov
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:10320-10340, 2026.

Abstract

Applying machine learning techniques to solving long-standing mathematical conjectures can be particularly challenging due to their extreme reward sparsity. As an illustrative example, we consider Kalai’s algebraic Hirsch conjecture and recast the construction of its counterexamples as a sparse-reward reinforcement learning problem on graphs. We propose a constrained options-based HRL framework with an equivariant graph neural network policy, which allows us to learn useful temporal abstractions for this task. We evaluate our approach over a wide range of degrees and demonstrate that it consistently outperforms classical RL algorithms as well as greedy search. By exploiting the hierarchical structure of the problem, we effectively provide a first-of-its-kind application of HRL to a problem in commutative algebra.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-butbaia26a, title = {Hierarchical Reinforcement Learning for Sparse-Reward Search in Commutative Algebra}, author = {Butbaia, Giorgi and Orland, Paul and Huang, Coco and Passaro, Davide and Fagan, Lucas and Tarquini, Michele and Dao, Hailong and Eisenbud, David and Shehper, Ali and Gukov, Sergei}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {10320--10340}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/butbaia26a/butbaia26a.pdf}, url = {https://proceedings.mlr.press/v306/butbaia26a.html}, abstract = {Applying machine learning techniques to solving long-standing mathematical conjectures can be particularly challenging due to their extreme reward sparsity. As an illustrative example, we consider Kalai’s algebraic Hirsch conjecture and recast the construction of its counterexamples as a sparse-reward reinforcement learning problem on graphs. We propose a constrained options-based HRL framework with an equivariant graph neural network policy, which allows us to learn useful temporal abstractions for this task. We evaluate our approach over a wide range of degrees and demonstrate that it consistently outperforms classical RL algorithms as well as greedy search. By exploiting the hierarchical structure of the problem, we effectively provide a first-of-its-kind application of HRL to a problem in commutative algebra.} }
Endnote
%0 Conference Paper %T Hierarchical Reinforcement Learning for Sparse-Reward Search in Commutative Algebra %A Giorgi Butbaia %A Paul Orland %A Coco Huang %A Davide Passaro %A Lucas Fagan %A Michele Tarquini %A Hailong Dao %A David Eisenbud %A Ali Shehper %A Sergei Gukov %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-butbaia26a %I PMLR %P 10320--10340 %U https://proceedings.mlr.press/v306/butbaia26a.html %V 306 %X Applying machine learning techniques to solving long-standing mathematical conjectures can be particularly challenging due to their extreme reward sparsity. As an illustrative example, we consider Kalai’s algebraic Hirsch conjecture and recast the construction of its counterexamples as a sparse-reward reinforcement learning problem on graphs. We propose a constrained options-based HRL framework with an equivariant graph neural network policy, which allows us to learn useful temporal abstractions for this task. We evaluate our approach over a wide range of degrees and demonstrate that it consistently outperforms classical RL algorithms as well as greedy search. By exploiting the hierarchical structure of the problem, we effectively provide a first-of-its-kind application of HRL to a problem in commutative algebra.
APA
Butbaia, G., Orland, P., Huang, C., Passaro, D., Fagan, L., Tarquini, M., Dao, H., Eisenbud, D., Shehper, A. & Gukov, S.. (2026). Hierarchical Reinforcement Learning for Sparse-Reward Search in Commutative Algebra. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:10320-10340 Available from https://proceedings.mlr.press/v306/butbaia26a.html.

Related Material