L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information Scaling

Zhuo Chen, Oriol Mayné I Comas, Zhuotao Jin, Di Luo, Marin Soljacic
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:13870-13909, 2026.

Abstract

Evaluating long-context language models on natural language conflates architectural capacity to capture dependencies with semantic knowledge and vocabulary statistics. When models fail at long contexts, we cannot determine whether failures stem from fundamental architectural limitations or insufficient domain knowledge, preventing clean diagnosis of efficient architectures before expensive training on real data. We introduce L-CUBE (Long-Context Utilization Benchmark), a synthetic benchmark that isolates dependency-capturing capacity from semantic knowledge through hierarchical Gaussian sequences with controllable bipartite mutual information scaling. The generator provides exact ground-truth conditionals that scale efficiently to arbitrarily long sequences, enabling unconfounded evaluation via conditional KL divergence rather than perplexity alone. We define long-context utilization to measure the amount of available predictive information that models extract as context grows. Experiments across transformers, state space models, and efficient alternatives validate L$^2$M capacity theory predictions and uncover new phenomena. L-CUBE enables practitioners to test whether a particular design will maintain long-context capability at target sequence lengths before committing to real-data training.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26q, title = {L-{CUBE}: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information Scaling}, author = {Chen, Zhuo and Comas, Oriol Mayn\'{e} I and Jin, Zhuotao and Luo, Di and Soljacic, Marin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {13870--13909}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26q/chen26q.pdf}, url = {https://proceedings.mlr.press/v306/chen26q.html}, abstract = {Evaluating long-context language models on natural language conflates architectural capacity to capture dependencies with semantic knowledge and vocabulary statistics. When models fail at long contexts, we cannot determine whether failures stem from fundamental architectural limitations or insufficient domain knowledge, preventing clean diagnosis of efficient architectures before expensive training on real data. We introduce L-CUBE (Long-Context Utilization Benchmark), a synthetic benchmark that isolates dependency-capturing capacity from semantic knowledge through hierarchical Gaussian sequences with controllable bipartite mutual information scaling. The generator provides exact ground-truth conditionals that scale efficiently to arbitrarily long sequences, enabling unconfounded evaluation via conditional KL divergence rather than perplexity alone. We define long-context utilization to measure the amount of available predictive information that models extract as context grows. Experiments across transformers, state space models, and efficient alternatives validate L$^2$M capacity theory predictions and uncover new phenomena. L-CUBE enables practitioners to test whether a particular design will maintain long-context capability at target sequence lengths before committing to real-data training.} }
Endnote
%0 Conference Paper %T L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information Scaling %A Zhuo Chen %A Oriol Mayné I Comas %A Zhuotao Jin %A Di Luo %A Marin Soljacic %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26q %I PMLR %P 13870--13909 %U https://proceedings.mlr.press/v306/chen26q.html %V 306 %X Evaluating long-context language models on natural language conflates architectural capacity to capture dependencies with semantic knowledge and vocabulary statistics. When models fail at long contexts, we cannot determine whether failures stem from fundamental architectural limitations or insufficient domain knowledge, preventing clean diagnosis of efficient architectures before expensive training on real data. We introduce L-CUBE (Long-Context Utilization Benchmark), a synthetic benchmark that isolates dependency-capturing capacity from semantic knowledge through hierarchical Gaussian sequences with controllable bipartite mutual information scaling. The generator provides exact ground-truth conditionals that scale efficiently to arbitrarily long sequences, enabling unconfounded evaluation via conditional KL divergence rather than perplexity alone. We define long-context utilization to measure the amount of available predictive information that models extract as context grows. Experiments across transformers, state space models, and efficient alternatives validate L$^2$M capacity theory predictions and uncover new phenomena. L-CUBE enables practitioners to test whether a particular design will maintain long-context capability at target sequence lengths before committing to real-data training.
APA
Chen, Z., Comas, O.M.I., Jin, Z., Luo, D. & Soljacic, M.. (2026). L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information Scaling. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:13870-13909 Available from https://proceedings.mlr.press/v306/chen26q.html.

Related Material