Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression

Xingwu Chen, Miao Lu, Beining Wu, Difan Zou
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:15885-15918, 2026.

Abstract

Scaling test-time computation during language model inference, such as generating intermediate thoughts or sampling multiple candidate answers, has proven effective in improving model performance. While these techniques inherently rely on the stochastic nature of inference to explore diverse reasoning paths, prior theoretical works typically build on a deterministic decoding framework, overlooking the stochastic nature of practical language model inference. This work takes an initial step to bridge this gap by establishing a new theoretical framework, incorporating randomness and sampling directly into the decoding analysis. To demonstrate the framework’s effectiveness, we apply it to the canonical in-context linear regression task with continuous and binary coefficients, simulating decoding via noise injection and sampling to analyze widely adopted inference techniques. We validate our theoretical findings through numerical simulations, with additional experiments on real-world tasks substantiating the framework’s potential for practical applications.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26cs, title = {Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression}, author = {Chen, Xingwu and Lu, Miao and Wu, Beining and Zou, Difan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {15885--15918}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26cs/chen26cs.pdf}, url = {https://proceedings.mlr.press/v306/chen26cs.html}, abstract = {Scaling test-time computation during language model inference, such as generating intermediate thoughts or sampling multiple candidate answers, has proven effective in improving model performance. While these techniques inherently rely on the stochastic nature of inference to explore diverse reasoning paths, prior theoretical works typically build on a deterministic decoding framework, overlooking the stochastic nature of practical language model inference. This work takes an initial step to bridge this gap by establishing a new theoretical framework, incorporating randomness and sampling directly into the decoding analysis. To demonstrate the framework’s effectiveness, we apply it to the canonical in-context linear regression task with continuous and binary coefficients, simulating decoding via noise injection and sampling to analyze widely adopted inference techniques. We validate our theoretical findings through numerical simulations, with additional experiments on real-world tasks substantiating the framework’s potential for practical applications.} }
Endnote
%0 Conference Paper %T Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression %A Xingwu Chen %A Miao Lu %A Beining Wu %A Difan Zou %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26cs %I PMLR %P 15885--15918 %U https://proceedings.mlr.press/v306/chen26cs.html %V 306 %X Scaling test-time computation during language model inference, such as generating intermediate thoughts or sampling multiple candidate answers, has proven effective in improving model performance. While these techniques inherently rely on the stochastic nature of inference to explore diverse reasoning paths, prior theoretical works typically build on a deterministic decoding framework, overlooking the stochastic nature of practical language model inference. This work takes an initial step to bridge this gap by establishing a new theoretical framework, incorporating randomness and sampling directly into the decoding analysis. To demonstrate the framework’s effectiveness, we apply it to the canonical in-context linear regression task with continuous and binary coefficients, simulating decoding via noise injection and sampling to analyze widely adopted inference techniques. We validate our theoretical findings through numerical simulations, with additional experiments on real-world tasks substantiating the framework’s potential for practical applications.
APA
Chen, X., Lu, M., Wu, B. & Zou, D.. (2026). Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:15885-15918 Available from https://proceedings.mlr.press/v306/chen26cs.html.

Related Material