Bayesian Symbolic Regression with Entropic Reinforcement Learning

Oussama Boussif, Mohammed Mahfoud, Younesse Kaddar, Moksh Jain, Sida Li, Damiano Fornasiere, Xiaoyin Chen, Yoshua Bengio, Esmeralda S. Whitammer
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:745-764, 2026.

Abstract

Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best-fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a {Bayesian} perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose {ERRLESS} (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum-entropy reinforcement learning. {ERRLESS} learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that {ERRLESS} achieves competitive results on the {Feynman} benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by {ERRLESS} achieves a high coefficient of determination ($R^2$) compared to an SMC baseline, highlighting the benefits of the {Bayesian} perspective in symbolic regression.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-boussif26a, title = {{Bayesian} Symbolic Regression with Entropic Reinforcement Learning}, author = {Boussif, Oussama and Mahfoud, Mohammed and Kaddar, Younesse and Jain, Moksh and Li, Sida and Fornasiere, Damiano and Chen, Xiaoyin and Bengio, Yoshua and Whitammer, Esmeralda S.}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {745--764}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/boussif26a/boussif26a.pdf}, url = {https://proceedings.mlr.press/v337/boussif26a.html}, abstract = {Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best-fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a {Bayesian} perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose {ERRLESS} (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum-entropy reinforcement learning. {ERRLESS} learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that {ERRLESS} achieves competitive results on the {Feynman} benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by {ERRLESS} achieves a high coefficient of determination ($R^2$) compared to an SMC baseline, highlighting the benefits of the {Bayesian} perspective in symbolic regression.} }
Endnote
%0 Conference Paper %T Bayesian Symbolic Regression with Entropic Reinforcement Learning %A Oussama Boussif %A Mohammed Mahfoud %A Younesse Kaddar %A Moksh Jain %A Sida Li %A Damiano Fornasiere %A Xiaoyin Chen %A Yoshua Bengio %A Esmeralda S. Whitammer %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-boussif26a %I PMLR %P 745--764 %U https://proceedings.mlr.press/v337/boussif26a.html %V 337 %X Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best-fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a {Bayesian} perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose {ERRLESS} (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum-entropy reinforcement learning. {ERRLESS} learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that {ERRLESS} achieves competitive results on the {Feynman} benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by {ERRLESS} achieves a high coefficient of determination ($R^2$) compared to an SMC baseline, highlighting the benefits of the {Bayesian} perspective in symbolic regression.
APA
Boussif, O., Mahfoud, M., Kaddar, Y., Jain, M., Li, S., Fornasiere, D., Chen, X., Bengio, Y. & Whitammer, E.S.. (2026). Bayesian Symbolic Regression with Entropic Reinforcement Learning. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:745-764 Available from https://proceedings.mlr.press/v337/boussif26a.html.

Related Material