Deriving Neural Scaling Laws from the Statistics of Natural Language

Francesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu Wyart
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:10450-10480, 2026.

Abstract

Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cagnetta26a, title = {Deriving Neural Scaling Laws from the Statistics of Natural Language}, author = {Cagnetta, Francesco and Raventos, Allan and Ganguli, Surya and Wyart, Matthieu}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {10450--10480}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cagnetta26a/cagnetta26a.pdf}, url = {https://proceedings.mlr.press/v306/cagnetta26a.html}, abstract = {Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.} }
Endnote
%0 Conference Paper %T Deriving Neural Scaling Laws from the Statistics of Natural Language %A Francesco Cagnetta %A Allan Raventos %A Surya Ganguli %A Matthieu Wyart %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cagnetta26a %I PMLR %P 10450--10480 %U https://proceedings.mlr.press/v306/cagnetta26a.html %V 306 %X Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.
APA
Cagnetta, F., Raventos, A., Ganguli, S. & Wyart, M.. (2026). Deriving Neural Scaling Laws from the Statistics of Natural Language. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:10450-10480 Available from https://proceedings.mlr.press/v306/cagnetta26a.html.

Related Material