Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization

Tina Behnia, Puneesh Deora, Christos Thrampoulidis
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:7314-7341, 2026.

Abstract

Language models are pretrained on sequences that blend statistical regularities, which make text fluent, with factual associations between specific tokens, which encode knowledge of facts. While recent works suggest that generalization depends critically on the interaction of these two streams, such as the diversity of the contexts in which facts appear, these effects remain difficult to study systematically. This paper introduces a flexible synthetic testbed that combines a statistical stream of generic tokens with an abstract factual stream of source-target token pairs, enabling fine-grained control over their interaction, such as their composition into a sequence (contextual structure) or the level of context diversity carrying the facts at training time. Through controlled experiments, we find that higher contextual diversity delays in-distribution factual learning, while low diversity can harm out-of-distribution generalization in ways that depend on the contextual structure. As a result, the optimal diversity level depends on the training budget. Beyond factual recall failures, we also identify failures in statistical generalization as we study how the interplay between contextual design and diversity level impacts different aspects of generalization. Furthermore, through a series of controlled interventions on the model components, we trace failure in different aspects of generalization to distinct optimization bottlenecks, highlighting the importance of the embedding and unembedding layers. Overall, our synthetic framework allows us to isolate effects that would be confounded in large-scale studies, offering a controlled testbed for future investigations.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-behnia26a, title = {Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization}, author = {Behnia, Tina and Deora, Puneesh and Thrampoulidis, Christos}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {7314--7341}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/behnia26a/behnia26a.pdf}, url = {https://proceedings.mlr.press/v306/behnia26a.html}, abstract = {Language models are pretrained on sequences that blend statistical regularities, which make text fluent, with factual associations between specific tokens, which encode knowledge of facts. While recent works suggest that generalization depends critically on the interaction of these two streams, such as the diversity of the contexts in which facts appear, these effects remain difficult to study systematically. This paper introduces a flexible synthetic testbed that combines a statistical stream of generic tokens with an abstract factual stream of source-target token pairs, enabling fine-grained control over their interaction, such as their composition into a sequence (contextual structure) or the level of context diversity carrying the facts at training time. Through controlled experiments, we find that higher contextual diversity delays in-distribution factual learning, while low diversity can harm out-of-distribution generalization in ways that depend on the contextual structure. As a result, the optimal diversity level depends on the training budget. Beyond factual recall failures, we also identify failures in statistical generalization as we study how the interplay between contextual design and diversity level impacts different aspects of generalization. Furthermore, through a series of controlled interventions on the model components, we trace failure in different aspects of generalization to distinct optimization bottlenecks, highlighting the importance of the embedding and unembedding layers. Overall, our synthetic framework allows us to isolate effects that would be confounded in large-scale studies, offering a controlled testbed for future investigations.} }
Endnote
%0 Conference Paper %T Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization %A Tina Behnia %A Puneesh Deora %A Christos Thrampoulidis %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-behnia26a %I PMLR %P 7314--7341 %U https://proceedings.mlr.press/v306/behnia26a.html %V 306 %X Language models are pretrained on sequences that blend statistical regularities, which make text fluent, with factual associations between specific tokens, which encode knowledge of facts. While recent works suggest that generalization depends critically on the interaction of these two streams, such as the diversity of the contexts in which facts appear, these effects remain difficult to study systematically. This paper introduces a flexible synthetic testbed that combines a statistical stream of generic tokens with an abstract factual stream of source-target token pairs, enabling fine-grained control over their interaction, such as their composition into a sequence (contextual structure) or the level of context diversity carrying the facts at training time. Through controlled experiments, we find that higher contextual diversity delays in-distribution factual learning, while low diversity can harm out-of-distribution generalization in ways that depend on the contextual structure. As a result, the optimal diversity level depends on the training budget. Beyond factual recall failures, we also identify failures in statistical generalization as we study how the interplay between contextual design and diversity level impacts different aspects of generalization. Furthermore, through a series of controlled interventions on the model components, we trace failure in different aspects of generalization to distinct optimization bottlenecks, highlighting the importance of the embedding and unembedding layers. Overall, our synthetic framework allows us to isolate effects that would be confounded in large-scale studies, offering a controlled testbed for future investigations.
APA
Behnia, T., Deora, P. & Thrampoulidis, C.. (2026). Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:7314-7341 Available from https://proceedings.mlr.press/v306/behnia26a.html.

Related Material