Uncovering the Gradient Geometry of Long CoT: A Spectral-guided Approach to Reasoning Distillation

Sinan Fan, Xiaofeng Sun, Chen Shen, Chenxi Huang, Shaotian Yan, Bing Wang, Kaiyuan Liu, Xiaosong Yuan, Liang Xie, Wenxiao Wang, Jun Zhang, Hongyang Chen, Jieping Ye
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:28824-28837, 2026.

Abstract

Large reasoning models (LRMs) achieve remarkable reasoning performance by generating long chains-of-thought (CoT). However, standard supervised fine-tuning (SFT) treats all tokens uniformly, indiscriminately minimizing loss across both essential reasoning steps and those that are noisy, redundant, or instance-specific. This often leads student models to memorize superficial patterns rather than acquire generalizable reasoning capabilities. To better understand this limitation, we introduce Loss Subspace Attribution, a gradient decomposition analysis approach that uncovers a striking geometric structure: Gradients corresponding to effective reasoning predominantly lie within a low-rank consensus subspace, while conflicting or unstructured signals dominate the residual subspace. Guided by this insight, we propose Spectral-guided Learning, a step-level distillation strategy that uses spectral strength to identify reasoning steps aligned with the consensus subspace and prioritizes their contribution to parameter updates, while suppressing gradients from the residual subspace. Experiments across various LRMs and diverse complex reasoning tasks consistently demonstrate that focusing optimization on the consensus subspace yields more robust and generalizable student models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fan26e, title = {Uncovering the Gradient Geometry of Long {C}o{T}: A Spectral-guided Approach to Reasoning Distillation}, author = {Fan, Sinan and Sun, Xiaofeng and Shen, Chen and Huang, Chenxi and Yan, Shaotian and Wang, Bing and Liu, Kaiyuan and Yuan, Xiaosong and Xie, Liang and Wang, Wenxiao and Zhang, Jun and Chen, Hongyang and Ye, Jieping}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {28824--28837}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fan26e/fan26e.pdf}, url = {https://proceedings.mlr.press/v306/fan26e.html}, abstract = {Large reasoning models (LRMs) achieve remarkable reasoning performance by generating long chains-of-thought (CoT). However, standard supervised fine-tuning (SFT) treats all tokens uniformly, indiscriminately minimizing loss across both essential reasoning steps and those that are noisy, redundant, or instance-specific. This often leads student models to memorize superficial patterns rather than acquire generalizable reasoning capabilities. To better understand this limitation, we introduce Loss Subspace Attribution, a gradient decomposition analysis approach that uncovers a striking geometric structure: Gradients corresponding to effective reasoning predominantly lie within a low-rank consensus subspace, while conflicting or unstructured signals dominate the residual subspace. Guided by this insight, we propose Spectral-guided Learning, a step-level distillation strategy that uses spectral strength to identify reasoning steps aligned with the consensus subspace and prioritizes their contribution to parameter updates, while suppressing gradients from the residual subspace. Experiments across various LRMs and diverse complex reasoning tasks consistently demonstrate that focusing optimization on the consensus subspace yields more robust and generalizable student models.} }
Endnote
%0 Conference Paper %T Uncovering the Gradient Geometry of Long CoT: A Spectral-guided Approach to Reasoning Distillation %A Sinan Fan %A Xiaofeng Sun %A Chen Shen %A Chenxi Huang %A Shaotian Yan %A Bing Wang %A Kaiyuan Liu %A Xiaosong Yuan %A Liang Xie %A Wenxiao Wang %A Jun Zhang %A Hongyang Chen %A Jieping Ye %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fan26e %I PMLR %P 28824--28837 %U https://proceedings.mlr.press/v306/fan26e.html %V 306 %X Large reasoning models (LRMs) achieve remarkable reasoning performance by generating long chains-of-thought (CoT). However, standard supervised fine-tuning (SFT) treats all tokens uniformly, indiscriminately minimizing loss across both essential reasoning steps and those that are noisy, redundant, or instance-specific. This often leads student models to memorize superficial patterns rather than acquire generalizable reasoning capabilities. To better understand this limitation, we introduce Loss Subspace Attribution, a gradient decomposition analysis approach that uncovers a striking geometric structure: Gradients corresponding to effective reasoning predominantly lie within a low-rank consensus subspace, while conflicting or unstructured signals dominate the residual subspace. Guided by this insight, we propose Spectral-guided Learning, a step-level distillation strategy that uses spectral strength to identify reasoning steps aligned with the consensus subspace and prioritizes their contribution to parameter updates, while suppressing gradients from the residual subspace. Experiments across various LRMs and diverse complex reasoning tasks consistently demonstrate that focusing optimization on the consensus subspace yields more robust and generalizable student models.
APA
Fan, S., Sun, X., Shen, C., Huang, C., Yan, S., Wang, B., Liu, K., Yuan, X., Xie, L., Wang, W., Zhang, J., Chen, H. & Ye, J.. (2026). Uncovering the Gradient Geometry of Long CoT: A Spectral-guided Approach to Reasoning Distillation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:28824-28837 Available from https://proceedings.mlr.press/v306/fan26e.html.

Related Material