Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization

Jiaxuan Cheng
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:18966-18978, 2026.

Abstract

Plasticity—a neural network’s ability to adapt to new tasks—is critical for continual and transfer learning. Existing measures, such as effective rank, dead neuron fraction, and weight norm, lack theoretical grounding and correlate poorly with performance on new tasks. We introduce local redundancy, an information-theoretic measure derived from universal compression theory. We define local redundancy as the worst-case redundancy of a local model family—parameters in an infinitesimal neighborhood along gradient directions—and show this is a principled measure of plasticity. Although local redundancy is intractable to compute exactly, we prove that the expected squared gradient norm on a synthetic memorization task provides an efficiently computable lower bound. Experiments on continual image classification and time series transfer learning demonstrate that local redundancy predicts downstream performance better than existing measures and enables pretraining checkpoint selection where validation loss plateaus.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cheng26i, title = {Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization}, author = {Cheng, Jiaxuan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {18966--18978}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cheng26i/cheng26i.pdf}, url = {https://proceedings.mlr.press/v306/cheng26i.html}, abstract = {Plasticity—a neural network’s ability to adapt to new tasks—is critical for continual and transfer learning. Existing measures, such as effective rank, dead neuron fraction, and weight norm, lack theoretical grounding and correlate poorly with performance on new tasks. We introduce local redundancy, an information-theoretic measure derived from universal compression theory. We define local redundancy as the worst-case redundancy of a local model family—parameters in an infinitesimal neighborhood along gradient directions—and show this is a principled measure of plasticity. Although local redundancy is intractable to compute exactly, we prove that the expected squared gradient norm on a synthetic memorization task provides an efficiently computable lower bound. Experiments on continual image classification and time series transfer learning demonstrate that local redundancy predicts downstream performance better than existing measures and enables pretraining checkpoint selection where validation loss plateaus.} }
Endnote
%0 Conference Paper %T Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization %A Jiaxuan Cheng %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cheng26i %I PMLR %P 18966--18978 %U https://proceedings.mlr.press/v306/cheng26i.html %V 306 %X Plasticity—a neural network’s ability to adapt to new tasks—is critical for continual and transfer learning. Existing measures, such as effective rank, dead neuron fraction, and weight norm, lack theoretical grounding and correlate poorly with performance on new tasks. We introduce local redundancy, an information-theoretic measure derived from universal compression theory. We define local redundancy as the worst-case redundancy of a local model family—parameters in an infinitesimal neighborhood along gradient directions—and show this is a principled measure of plasticity. Although local redundancy is intractable to compute exactly, we prove that the expected squared gradient norm on a synthetic memorization task provides an efficiently computable lower bound. Experiments on continual image classification and time series transfer learning demonstrate that local redundancy predicts downstream performance better than existing measures and enables pretraining checkpoint selection where validation loss plateaus.
APA
Cheng, J.. (2026). Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:18966-18978 Available from https://proceedings.mlr.press/v306/cheng26i.html.

Related Material