Bits That Count: Quantifying and Predicting Capabilities of Language Models

Elizabeth Donoway, Hailey Joren, Michael R Deweese, Ethan Perez, John Schulman, Fabien Roger, Jan Leike
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:26127-26151, 2026.

Abstract

When does learning elicit existing knowledge, and when does it primarily teach new capabilities? We find that the amount of generalizable information language models learn during training predicts the origins of their emergent capabilities. Minuscule amounts of information—in many cases, a few bits in a single example—can unlock large fractions of models’ maximum performance when capabilities are elicited rather than taught. We quantify these learning regimes using excess description length (EDL), an information-theoretic measure of generalizable information learned during training. We find that elicitation and teaching exhibit distinct EDL signatures that characterize the predominant learning mechanism as information scales: elicitation requires orders of magnitude less information than teaching to comparable performance. We demonstrate that EDL provides a practical tool for quantitatively estimating the maximum amount of predictive information models can compress from data into trainable parameters during learning. These capacity limits describe optimal tradeoffs between data and parameter count that robustly predict when parameter-efficient fine-tuning methods (e.g., LoRA) will underperform full fine-tuning.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-donoway26a, title = {Bits That Count: Quantifying and Predicting Capabilities of Language Models}, author = {Donoway, Elizabeth and Joren, Hailey and Deweese, Michael R and Perez, Ethan and Schulman, John and Roger, Fabien and Leike, Jan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {26127--26151}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/donoway26a/donoway26a.pdf}, url = {https://proceedings.mlr.press/v306/donoway26a.html}, abstract = {When does learning elicit existing knowledge, and when does it primarily teach new capabilities? We find that the amount of generalizable information language models learn during training predicts the origins of their emergent capabilities. Minuscule amounts of information—in many cases, a few bits in a single example—can unlock large fractions of models’ maximum performance when capabilities are elicited rather than taught. We quantify these learning regimes using excess description length (EDL), an information-theoretic measure of generalizable information learned during training. We find that elicitation and teaching exhibit distinct EDL signatures that characterize the predominant learning mechanism as information scales: elicitation requires orders of magnitude less information than teaching to comparable performance. We demonstrate that EDL provides a practical tool for quantitatively estimating the maximum amount of predictive information models can compress from data into trainable parameters during learning. These capacity limits describe optimal tradeoffs between data and parameter count that robustly predict when parameter-efficient fine-tuning methods (e.g., LoRA) will underperform full fine-tuning.} }
Endnote
%0 Conference Paper %T Bits That Count: Quantifying and Predicting Capabilities of Language Models %A Elizabeth Donoway %A Hailey Joren %A Michael R Deweese %A Ethan Perez %A John Schulman %A Fabien Roger %A Jan Leike %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-donoway26a %I PMLR %P 26127--26151 %U https://proceedings.mlr.press/v306/donoway26a.html %V 306 %X When does learning elicit existing knowledge, and when does it primarily teach new capabilities? We find that the amount of generalizable information language models learn during training predicts the origins of their emergent capabilities. Minuscule amounts of information—in many cases, a few bits in a single example—can unlock large fractions of models’ maximum performance when capabilities are elicited rather than taught. We quantify these learning regimes using excess description length (EDL), an information-theoretic measure of generalizable information learned during training. We find that elicitation and teaching exhibit distinct EDL signatures that characterize the predominant learning mechanism as information scales: elicitation requires orders of magnitude less information than teaching to comparable performance. We demonstrate that EDL provides a practical tool for quantitatively estimating the maximum amount of predictive information models can compress from data into trainable parameters during learning. These capacity limits describe optimal tradeoffs between data and parameter count that robustly predict when parameter-efficient fine-tuning methods (e.g., LoRA) will underperform full fine-tuning.
APA
Donoway, E., Joren, H., Deweese, M.R., Perez, E., Schulman, J., Roger, F. & Leike, J.. (2026). Bits That Count: Quantifying and Predicting Capabilities of Language Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:26127-26151 Available from https://proceedings.mlr.press/v306/donoway26a.html.

Related Material