Rank Lifting and Random Non-Linear Maps

Andrea Drago, Maria Sofia Bucarelli, Francesco Caso, Marius Michetti, Federico Siciliano, Fabrizio Silvestri, Luca Becchetti
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2881-2889, 2026.

Abstract

Deep neural networks exhibit improved training and generalization performance as the number of parameters grows well beyond the size of the training set, contradicting classical intuitions about overfitting. In order to gain a better understanding of this “benign overparameterization”, we analyze the representational capacity of a random one-hidden-layer perceptron with Gaussian weights, no bias and threshold activations. More precisely, we investigate the following question: when does a hidden layer of dimension $n$ maps $k$ input vectors with pairwise angles at least $\theta$, to a full-rank activation matrix, thus ensuring that a simple linear classifier can perfectly fit those inputs in feature space? This problem has an immediate impact on memorization capacity at initialization and we frame it as a question about hyperplane arrangements on the unit sphere, and we prove new isoperimetric-like inequalities. This allows us to derive non-trivial lower bounds on the probability that a random embedding avoids the arrangement’s zero-measure regions. Our results show that once the hidden dimension exceeds a threshold (depending on $\theta$ and the input dimension), hidden representations are linearly independent with high probability. While the case we consider is challenging due to the sparsity of the solution space, this setting highlights crucial, underlying geometric problems and connections to related questions in spherical geometry and linear algebra.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-drago26a, title = { Rank Lifting and Random Non-Linear Maps }, author = {Drago, Andrea and Bucarelli, Maria Sofia and Caso, Francesco and Michetti, Marius and Siciliano, Federico and Silvestri, Fabrizio and Becchetti, Luca}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2881--2889}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/drago26a/drago26a.pdf}, url = {https://proceedings.mlr.press/v300/drago26a.html}, abstract = { Deep neural networks exhibit improved training and generalization performance as the number of parameters grows well beyond the size of the training set, contradicting classical intuitions about overfitting. In order to gain a better understanding of this “benign overparameterization”, we analyze the representational capacity of a random one-hidden-layer perceptron with Gaussian weights, no bias and threshold activations. More precisely, we investigate the following question: when does a hidden layer of dimension $n$ maps $k$ input vectors with pairwise angles at least $\theta$, to a full-rank activation matrix, thus ensuring that a simple linear classifier can perfectly fit those inputs in feature space? This problem has an immediate impact on memorization capacity at initialization and we frame it as a question about hyperplane arrangements on the unit sphere, and we prove new isoperimetric-like inequalities. This allows us to derive non-trivial lower bounds on the probability that a random embedding avoids the arrangement’s zero-measure regions. Our results show that once the hidden dimension exceeds a threshold (depending on $\theta$ and the input dimension), hidden representations are linearly independent with high probability. While the case we consider is challenging due to the sparsity of the solution space, this setting highlights crucial, underlying geometric problems and connections to related questions in spherical geometry and linear algebra. } }
Endnote
%0 Conference Paper %T Rank Lifting and Random Non-Linear Maps %A Andrea Drago %A Maria Sofia Bucarelli %A Francesco Caso %A Marius Michetti %A Federico Siciliano %A Fabrizio Silvestri %A Luca Becchetti %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-drago26a %I PMLR %P 2881--2889 %U https://proceedings.mlr.press/v300/drago26a.html %V 300 %X Deep neural networks exhibit improved training and generalization performance as the number of parameters grows well beyond the size of the training set, contradicting classical intuitions about overfitting. In order to gain a better understanding of this “benign overparameterization”, we analyze the representational capacity of a random one-hidden-layer perceptron with Gaussian weights, no bias and threshold activations. More precisely, we investigate the following question: when does a hidden layer of dimension $n$ maps $k$ input vectors with pairwise angles at least $\theta$, to a full-rank activation matrix, thus ensuring that a simple linear classifier can perfectly fit those inputs in feature space? This problem has an immediate impact on memorization capacity at initialization and we frame it as a question about hyperplane arrangements on the unit sphere, and we prove new isoperimetric-like inequalities. This allows us to derive non-trivial lower bounds on the probability that a random embedding avoids the arrangement’s zero-measure regions. Our results show that once the hidden dimension exceeds a threshold (depending on $\theta$ and the input dimension), hidden representations are linearly independent with high probability. While the case we consider is challenging due to the sparsity of the solution space, this setting highlights crucial, underlying geometric problems and connections to related questions in spherical geometry and linear algebra.
APA
Drago, A., Bucarelli, M.S., Caso, F., Michetti, M., Siciliano, F., Silvestri, F. & Becchetti, L.. (2026). Rank Lifting and Random Non-Linear Maps . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2881-2889 Available from https://proceedings.mlr.press/v300/drago26a.html.

Related Material