Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers

Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma, William G. Underwood, Richard J. Samworth
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:19581-19610, 2026.

Abstract

We study in-context learning for nonparametric regression with $\alpha$-Hölder smooth regression functions, for some $\alpha>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained transformer with $\Theta(\log n)$ parameters and $\Omega(n^{2\alpha/(2\alpha+d)}\log^3 n)$ pretraining sequences can achieve the minimax optimal rate of convergence $O(n^{-2\alpha/(2\alpha+d)})$ in mean squared error. Our result requires substantially fewer transformer parameters and pretraining sequences than previous results in the literature. This is achieved by showing that transformers are able to approximate local polynomial estimators efficiently by implementing a kernel-weighted polynomial basis and then running gradient descent.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ching26a, title = {Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers}, author = {Ching, Michelle and Popescu, Ioana and Smith, Nico and Ma, Tianyi and Underwood, William G. and Samworth, Richard J.}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {19581--19610}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ching26a/ching26a.pdf}, url = {https://proceedings.mlr.press/v306/ching26a.html}, abstract = {We study in-context learning for nonparametric regression with $\alpha$-Hölder smooth regression functions, for some $\alpha>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained transformer with $\Theta(\log n)$ parameters and $\Omega(n^{2\alpha/(2\alpha+d)}\log^3 n)$ pretraining sequences can achieve the minimax optimal rate of convergence $O(n^{-2\alpha/(2\alpha+d)})$ in mean squared error. Our result requires substantially fewer transformer parameters and pretraining sequences than previous results in the literature. This is achieved by showing that transformers are able to approximate local polynomial estimators efficiently by implementing a kernel-weighted polynomial basis and then running gradient descent.} }
Endnote
%0 Conference Paper %T Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers %A Michelle Ching %A Ioana Popescu %A Nico Smith %A Tianyi Ma %A William G. Underwood %A Richard J. Samworth %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ching26a %I PMLR %P 19581--19610 %U https://proceedings.mlr.press/v306/ching26a.html %V 306 %X We study in-context learning for nonparametric regression with $\alpha$-Hölder smooth regression functions, for some $\alpha>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained transformer with $\Theta(\log n)$ parameters and $\Omega(n^{2\alpha/(2\alpha+d)}\log^3 n)$ pretraining sequences can achieve the minimax optimal rate of convergence $O(n^{-2\alpha/(2\alpha+d)})$ in mean squared error. Our result requires substantially fewer transformer parameters and pretraining sequences than previous results in the literature. This is achieved by showing that transformers are able to approximate local polynomial estimators efficiently by implementing a kernel-weighted polynomial basis and then running gradient descent.
APA
Ching, M., Popescu, I., Smith, N., Ma, T., Underwood, W.G. & Samworth, R.J.. (2026). Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:19581-19610 Available from https://proceedings.mlr.press/v306/ching26a.html.

Related Material