[edit]
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:26202-26226, 2026.
Abstract
Theoretical analyses of Empirical Risk Minimization (ERM) are standardly framed within the Real-RAM model of computation. In this setting, training even simple neural networks is known to be $\exists \mathbb{R}$-complete - a complexity class believed to be harder than NP, characterizing the difficulty of solving systems of polynomial inequalities over the real numbers. However, this algebraic framework diverges from the reality of digital computation with finite-precision hardware. In this work, we analyze the theoretical complexity of ERM under a realistic bit-level model (ERM-bit), where network parameters and inputs are constrained to be rational numbers with polynomially bounded bit-lengths. Under this model, we reveal a sharp dichotomy in tractability governed by the activation function: for deep networks with any polynomial activation with rational coefficients and degree at least $2$, deciding ERM-bit is #P-hard, determining the sign of a single partial derivative is unlikely to be in BPP, and deciding a specific bit in the gradient is #P-hard. In contrast, for piecewise-linear activations such as ReLU, precision requirements remain manageable: ERM-bit is in NP, indeed NP-complete, and standard backpropagation runs in polynomial time, showing that finite-precision constraints are not merely implementation details but fundamental determinants of learnability.