The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution

Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang, Haokun Chen, Guancheng Wan, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:8076-8099, 2026.

Abstract

Large Reasoning Models (LRMs) enhance performance by generating explicit Chain-of-Thought (CoT) trajectories, yet enabling them to self-evaluate correctness without external supervision remains a critical challenge. Existing methods often rely on ground-truth labels or shallow output probabilities, neglecting the layerwise evolution of the reasoning trajectory. In this work, we introduce GeoR (Geometry of Reasoning), a white-box self-evaluation framework based on layerwise trajectory evolution. GeoR decomposes reasoning fidelity into two complementary dimensions: (1) Geometric Evolution, which synthesizes the first- and second-order evolution of layerwise hidden-state trajectories to quantify geometric progress in reasoning; and (2) Difficulty-Aware Calibration, which utilizes cross-entropy of reasoning progress to normalize the Geometric Evolution against intrinsic query uncertainty. By jointly modeling these factors, GeoR effectively distinguishes the coherent evolution of correct reasoning from the chaotic trajectories of errors. Extensive experiments across eight LRMs and seven benchmarks demonstrate that GeoR consistently outperforms state-of-the-art baselines in AUROC, AUPR, and FPR@95.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bi26e, title = {The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution}, author = {Bi, Jinhe and Yan, Danqi and Wang, Yifan and Huang, Wenke and Chen, Haokun and Wan, Guancheng and Ye, Mang and Xiao, Xun and Schuetze, Hinrich and Tresp, Volker and Ma, Yunpu}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {8076--8099}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bi26e/bi26e.pdf}, url = {https://proceedings.mlr.press/v306/bi26e.html}, abstract = {Large Reasoning Models (LRMs) enhance performance by generating explicit Chain-of-Thought (CoT) trajectories, yet enabling them to self-evaluate correctness without external supervision remains a critical challenge. Existing methods often rely on ground-truth labels or shallow output probabilities, neglecting the layerwise evolution of the reasoning trajectory. In this work, we introduce GeoR (Geometry of Reasoning), a white-box self-evaluation framework based on layerwise trajectory evolution. GeoR decomposes reasoning fidelity into two complementary dimensions: (1) Geometric Evolution, which synthesizes the first- and second-order evolution of layerwise hidden-state trajectories to quantify geometric progress in reasoning; and (2) Difficulty-Aware Calibration, which utilizes cross-entropy of reasoning progress to normalize the Geometric Evolution against intrinsic query uncertainty. By jointly modeling these factors, GeoR effectively distinguishes the coherent evolution of correct reasoning from the chaotic trajectories of errors. Extensive experiments across eight LRMs and seven benchmarks demonstrate that GeoR consistently outperforms state-of-the-art baselines in AUROC, AUPR, and FPR@95.} }
Endnote
%0 Conference Paper %T The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution %A Jinhe Bi %A Danqi Yan %A Yifan Wang %A Wenke Huang %A Haokun Chen %A Guancheng Wan %A Mang Ye %A Xun Xiao %A Hinrich Schuetze %A Volker Tresp %A Yunpu Ma %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bi26e %I PMLR %P 8076--8099 %U https://proceedings.mlr.press/v306/bi26e.html %V 306 %X Large Reasoning Models (LRMs) enhance performance by generating explicit Chain-of-Thought (CoT) trajectories, yet enabling them to self-evaluate correctness without external supervision remains a critical challenge. Existing methods often rely on ground-truth labels or shallow output probabilities, neglecting the layerwise evolution of the reasoning trajectory. In this work, we introduce GeoR (Geometry of Reasoning), a white-box self-evaluation framework based on layerwise trajectory evolution. GeoR decomposes reasoning fidelity into two complementary dimensions: (1) Geometric Evolution, which synthesizes the first- and second-order evolution of layerwise hidden-state trajectories to quantify geometric progress in reasoning; and (2) Difficulty-Aware Calibration, which utilizes cross-entropy of reasoning progress to normalize the Geometric Evolution against intrinsic query uncertainty. By jointly modeling these factors, GeoR effectively distinguishes the coherent evolution of correct reasoning from the chaotic trajectories of errors. Extensive experiments across eight LRMs and seven benchmarks demonstrate that GeoR consistently outperforms state-of-the-art baselines in AUROC, AUPR, and FPR@95.
APA
Bi, J., Yan, D., Wang, Y., Huang, W., Chen, H., Wan, G., Ye, M., Xiao, X., Schuetze, H., Tresp, V. & Ma, Y.. (2026). The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:8076-8099 Available from https://proceedings.mlr.press/v306/bi26e.html.

Related Material