Reasoning Structure of Large Language Models

Frédéric Berdoz, Luca A Lanzendörfer, Fabian Farestam, Roger Wattenhofer
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:7613-7640, 2026.

Abstract

Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics can hide fundamentally different reasoning structures. To address this limitation, we introduce a scalable LRM benchmark of logic puzzles and a pipeline that converts unstructured traces into verifiable reasoning graphs of claims and dependencies. This turns reasoning into a structured, measurable object whose topology can be quantitatively analyzed. Building on this, we define a reasoning efficiency metric that quantifies how concentrated the model’s logical flow is. Our analysis on open-source reasoning models shows that structural measurements separate behaviors that token count and accuracy conflate, providing a practical tool for diagnosing failure modes and comparing how reasoning scales with puzzle difficulty.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-berdoz26b, title = {Reasoning Structure of Large Language Models}, author = {Berdoz, Fr\'{e}d\'{e}ric and Lanzend\"{o}rfer, Luca A and Farestam, Fabian and Wattenhofer, Roger}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {7613--7640}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/berdoz26b/berdoz26b.pdf}, url = {https://proceedings.mlr.press/v306/berdoz26b.html}, abstract = {Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics can hide fundamentally different reasoning structures. To address this limitation, we introduce a scalable LRM benchmark of logic puzzles and a pipeline that converts unstructured traces into verifiable reasoning graphs of claims and dependencies. This turns reasoning into a structured, measurable object whose topology can be quantitatively analyzed. Building on this, we define a reasoning efficiency metric that quantifies how concentrated the model’s logical flow is. Our analysis on open-source reasoning models shows that structural measurements separate behaviors that token count and accuracy conflate, providing a practical tool for diagnosing failure modes and comparing how reasoning scales with puzzle difficulty.} }
Endnote
%0 Conference Paper %T Reasoning Structure of Large Language Models %A Frédéric Berdoz %A Luca A Lanzendörfer %A Fabian Farestam %A Roger Wattenhofer %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-berdoz26b %I PMLR %P 7613--7640 %U https://proceedings.mlr.press/v306/berdoz26b.html %V 306 %X Large reasoning models (LRMs) are often evaluated using metrics such as final-answer accuracy or token count. However, identical scores on these metrics can hide fundamentally different reasoning structures. To address this limitation, we introduce a scalable LRM benchmark of logic puzzles and a pipeline that converts unstructured traces into verifiable reasoning graphs of claims and dependencies. This turns reasoning into a structured, measurable object whose topology can be quantitatively analyzed. Building on this, we define a reasoning efficiency metric that quantifies how concentrated the model’s logical flow is. Our analysis on open-source reasoning models shows that structural measurements separate behaviors that token count and accuracy conflate, providing a practical tool for diagnosing failure modes and comparing how reasoning scales with puzzle difficulty.
APA
Berdoz, F., Lanzendörfer, L.A., Farestam, F. & Wattenhofer, R.. (2026). Reasoning Structure of Large Language Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:7613-7640 Available from https://proceedings.mlr.press/v306/berdoz26b.html.

Related Material