Context-level Language Modeling by Learning Predictive Context Embeddings

Beiya Dai, Yuliang Liu, Yunchong Song, Daozheng Xue, Qipeng Guo, Kai Chen, Xinbing Wang, Bowen Zhou, Zhouhan Lin
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:22528-22545, 2026.

Abstract

We propose ContextLM, a framework that implicitly learns multi-token prediction by augmenting standard pretraining with an intrinsic next-context prediction objective. ContextLM builds a language model on top of context embeddings that span multiple tokens, enabling better next-token prediction by predicting the next context. Our model is fully compatible with standard autoregressive, token-by-token evaluation paradigms (e.g., perplexity). Extensive experiments with GPT-2 and Pythia backbones (up to 1.5B parameters and 300B training tokens) reveal that ContextLM shifts the Pareto frontier of scaling laws, exhibiting superior efficiency in parameters, training tokens, and FLOPs. Our results show that ContextLM could already achieve the baseline perplexity using 39% fewer parameters and demonstrates robust generalization improvements on extensive downstream tasks under equivalent parameter counts.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-dai26g, title = {Context-level Language Modeling by Learning Predictive Context Embeddings}, author = {Dai, Beiya and Liu, Yuliang and Song, Yunchong and Xue, Daozheng and Guo, Qipeng and Chen, Kai and Wang, Xinbing and Zhou, Bowen and Lin, Zhouhan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {22528--22545}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/dai26g/dai26g.pdf}, url = {https://proceedings.mlr.press/v306/dai26g.html}, abstract = {We propose ContextLM, a framework that implicitly learns multi-token prediction by augmenting standard pretraining with an intrinsic next-context prediction objective. ContextLM builds a language model on top of context embeddings that span multiple tokens, enabling better next-token prediction by predicting the next context. Our model is fully compatible with standard autoregressive, token-by-token evaluation paradigms (e.g., perplexity). Extensive experiments with GPT-2 and Pythia backbones (up to 1.5B parameters and 300B training tokens) reveal that ContextLM shifts the Pareto frontier of scaling laws, exhibiting superior efficiency in parameters, training tokens, and FLOPs. Our results show that ContextLM could already achieve the baseline perplexity using 39% fewer parameters and demonstrates robust generalization improvements on extensive downstream tasks under equivalent parameter counts.} }
Endnote
%0 Conference Paper %T Context-level Language Modeling by Learning Predictive Context Embeddings %A Beiya Dai %A Yuliang Liu %A Yunchong Song %A Daozheng Xue %A Qipeng Guo %A Kai Chen %A Xinbing Wang %A Bowen Zhou %A Zhouhan Lin %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-dai26g %I PMLR %P 22528--22545 %U https://proceedings.mlr.press/v306/dai26g.html %V 306 %X We propose ContextLM, a framework that implicitly learns multi-token prediction by augmenting standard pretraining with an intrinsic next-context prediction objective. ContextLM builds a language model on top of context embeddings that span multiple tokens, enabling better next-token prediction by predicting the next context. Our model is fully compatible with standard autoregressive, token-by-token evaluation paradigms (e.g., perplexity). Extensive experiments with GPT-2 and Pythia backbones (up to 1.5B parameters and 300B training tokens) reveal that ContextLM shifts the Pareto frontier of scaling laws, exhibiting superior efficiency in parameters, training tokens, and FLOPs. Our results show that ContextLM could already achieve the baseline perplexity using 39% fewer parameters and demonstrates robust generalization improvements on extensive downstream tasks under equivalent parameter counts.
APA
Dai, B., Liu, Y., Song, Y., Xue, D., Guo, Q., Chen, K., Wang, X., Zhou, B. & Lin, Z.. (2026). Context-level Language Modeling by Learning Predictive Context Embeddings. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:22528-22545 Available from https://proceedings.mlr.press/v306/dai26g.html.

Related Material