Chain-of-Thought Gradient Descent

Hong-Yu Chen, Venkat Sripad Ganti, Hude Liu, Jerry Yao-Chieh Hu, Han Liu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:14232-14278, 2026.

Abstract

We show that Chain-of-Thought (CoT) expands the expressiveness of Transformer in-context learning (ICL). Specifically, we show CoT enable efficient simulation of In-Context Gradient Descent (ICGD) for $N$-layer neural network. Different from CoT, a Transformer with fixed depth and hidden dimension has fixed ICL capacity in one forward pass. Simulating larger models or more optimization steps in-context requires deeper or wider Transformers. CoT removes this limitation by providing an expandable workspace via the sequence trajectory. This enables arbitrary-step and arbitrary-capacity ICGD within a constant-depth Transformer. Second, we provide a provable efficient guarantee unique to CoT through dynamical masking. The attention mechanism only process the relevant tokens for the current update step. This eliminates the redundant “process everything” cost of single-pass deep models. Specifically, we prove this CoT mechanism improves the computational cost of the prior best in-context result [Wu et al., ICML 2025] by $O(N)$. Numerical validations support our theory.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26ae, title = {Chain-of-Thought Gradient Descent}, author = {Chen, Hong-Yu and Ganti, Venkat Sripad and Liu, Hude and Hu, Jerry Yao-Chieh and Liu, Han}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {14232--14278}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26ae/chen26ae.pdf}, url = {https://proceedings.mlr.press/v306/chen26ae.html}, abstract = {We show that Chain-of-Thought (CoT) expands the expressiveness of Transformer in-context learning (ICL). Specifically, we show CoT enable efficient simulation of In-Context Gradient Descent (ICGD) for $N$-layer neural network. Different from CoT, a Transformer with fixed depth and hidden dimension has fixed ICL capacity in one forward pass. Simulating larger models or more optimization steps in-context requires deeper or wider Transformers. CoT removes this limitation by providing an expandable workspace via the sequence trajectory. This enables arbitrary-step and arbitrary-capacity ICGD within a constant-depth Transformer. Second, we provide a provable efficient guarantee unique to CoT through dynamical masking. The attention mechanism only process the relevant tokens for the current update step. This eliminates the redundant “process everything” cost of single-pass deep models. Specifically, we prove this CoT mechanism improves the computational cost of the prior best in-context result [Wu et al., ICML 2025] by $O(N)$. Numerical validations support our theory.} }
Endnote
%0 Conference Paper %T Chain-of-Thought Gradient Descent %A Hong-Yu Chen %A Venkat Sripad Ganti %A Hude Liu %A Jerry Yao-Chieh Hu %A Han Liu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26ae %I PMLR %P 14232--14278 %U https://proceedings.mlr.press/v306/chen26ae.html %V 306 %X We show that Chain-of-Thought (CoT) expands the expressiveness of Transformer in-context learning (ICL). Specifically, we show CoT enable efficient simulation of In-Context Gradient Descent (ICGD) for $N$-layer neural network. Different from CoT, a Transformer with fixed depth and hidden dimension has fixed ICL capacity in one forward pass. Simulating larger models or more optimization steps in-context requires deeper or wider Transformers. CoT removes this limitation by providing an expandable workspace via the sequence trajectory. This enables arbitrary-step and arbitrary-capacity ICGD within a constant-depth Transformer. Second, we provide a provable efficient guarantee unique to CoT through dynamical masking. The attention mechanism only process the relevant tokens for the current update step. This eliminates the redundant “process everything” cost of single-pass deep models. Specifically, we prove this CoT mechanism improves the computational cost of the prior best in-context result [Wu et al., ICML 2025] by $O(N)$. Numerical validations support our theory.
APA
Chen, H., Ganti, V.S., Liu, H., Hu, J.Y. & Liu, H.. (2026). Chain-of-Thought Gradient Descent. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:14232-14278 Available from https://proceedings.mlr.press/v306/chen26ae.html.

Related Material