Learning Linear Regression with Low-Rank Tasks In-Context

Kaito Takanami, Takashi Takahashi, Yoshiyuki Kabashima
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4672-4680, 2026.

Abstract

In-context learning (ICL) is a key building block of modern large language models, yet its theoretical mechanisms remain poorly understood. It is particularly mysterious how ICL operates in real-world applications where tasks have a structure. In this work, we address this problem by analyzing a linear attention model trained on low-rank regression tasks. Within this setting, we precisely characterize the distribution of predictions and the generalization error in the high-dimensional limit. Moreover, we find that statistical fluctuations in finite pre-training data induce an implicit regularization. Finally, we identify a sharp phase transition of the generalization error governed by task structure. These results provide a framework for understanding how transformers learn to learn the task structure.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-takanami26a, title = { Learning Linear Regression with Low-Rank Tasks In-Context }, author = {Takanami, Kaito and Takahashi, Takashi and Kabashima, Yoshiyuki}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4672--4680}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/takanami26a/takanami26a.pdf}, url = {https://proceedings.mlr.press/v300/takanami26a.html}, abstract = { In-context learning (ICL) is a key building block of modern large language models, yet its theoretical mechanisms remain poorly understood. It is particularly mysterious how ICL operates in real-world applications where tasks have a structure. In this work, we address this problem by analyzing a linear attention model trained on low-rank regression tasks. Within this setting, we precisely characterize the distribution of predictions and the generalization error in the high-dimensional limit. Moreover, we find that statistical fluctuations in finite pre-training data induce an implicit regularization. Finally, we identify a sharp phase transition of the generalization error governed by task structure. These results provide a framework for understanding how transformers learn to learn the task structure. } }
Endnote
%0 Conference Paper %T Learning Linear Regression with Low-Rank Tasks In-Context %A Kaito Takanami %A Takashi Takahashi %A Yoshiyuki Kabashima %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-takanami26a %I PMLR %P 4672--4680 %U https://proceedings.mlr.press/v300/takanami26a.html %V 300 %X In-context learning (ICL) is a key building block of modern large language models, yet its theoretical mechanisms remain poorly understood. It is particularly mysterious how ICL operates in real-world applications where tasks have a structure. In this work, we address this problem by analyzing a linear attention model trained on low-rank regression tasks. Within this setting, we precisely characterize the distribution of predictions and the generalization error in the high-dimensional limit. Moreover, we find that statistical fluctuations in finite pre-training data induce an implicit regularization. Finally, we identify a sharp phase transition of the generalization error governed by task structure. These results provide a framework for understanding how transformers learn to learn the task structure.
APA
Takanami, K., Takahashi, T. & Kabashima, Y.. (2026). Learning Linear Regression with Low-Rank Tasks In-Context . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4672-4680 Available from https://proceedings.mlr.press/v300/takanami26a.html.

Related Material