Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking

Jing Bi, Luchuan Song, Dingxin Zhang, Pinxin Liu, Guangyu Sun, Lianggong Bruce Wen, Weidong Cai, Chen Chen, Chenliang Xu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:8063-8075, 2026.

Abstract

Vision-language models (VLMs) now support both direct Instruct and explicit-reasoning Thinking modes, yet practitioners still lack principled ways to decide when reasoning actually improves performance, or how much computation to spend at test time, so we investigate whether VLMs encode meta-cognitive signals for adaptive inference. We derive oracle labels for two properties: (1) reasoning helpfulness, namely whether explicit reasoning improves accuracy, and (2) desired generation length, the minimal token budget needed for a correct answer. Probing final-layer representations in InternVL and Qwen3-VL models, we find Thinking models encode these signals more linearly than Instruct models, suggesting that reasoning-oriented training enhances meta-cognitive structure. Head-wise attribution reveals two circuits: length-control heads in lower layers and reasoning/difficulty heads in higher layers. Causal interventions confirm these roles across benchmarks: scaling length heads controls output length with minimal accuracy loss, while scaling reasoning heads enables a perception-reasoning trade-off, improving accuracy by up to 5.3%. Our results demonstrate that reasoning-tuned VLMs develop localized, manipulable circuits for meta-cognitive control, enabling test-time steering of computation and reasoning without retraining.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bi26d, title = {Does Reasoning Improve Seeing? {U}nderstanding When Vision-Language Models Benefit from Thinking}, author = {Bi, Jing and Song, Luchuan and Zhang, Dingxin and Liu, Pinxin and Sun, Guangyu and Wen, Lianggong Bruce and Cai, Weidong and Chen, Chen and Xu, Chenliang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {8063--8075}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bi26d/bi26d.pdf}, url = {https://proceedings.mlr.press/v306/bi26d.html}, abstract = {Vision-language models (VLMs) now support both direct Instruct and explicit-reasoning Thinking modes, yet practitioners still lack principled ways to decide when reasoning actually improves performance, or how much computation to spend at test time, so we investigate whether VLMs encode meta-cognitive signals for adaptive inference. We derive oracle labels for two properties: (1) reasoning helpfulness, namely whether explicit reasoning improves accuracy, and (2) desired generation length, the minimal token budget needed for a correct answer. Probing final-layer representations in InternVL and Qwen3-VL models, we find Thinking models encode these signals more linearly than Instruct models, suggesting that reasoning-oriented training enhances meta-cognitive structure. Head-wise attribution reveals two circuits: length-control heads in lower layers and reasoning/difficulty heads in higher layers. Causal interventions confirm these roles across benchmarks: scaling length heads controls output length with minimal accuracy loss, while scaling reasoning heads enables a perception-reasoning trade-off, improving accuracy by up to 5.3%. Our results demonstrate that reasoning-tuned VLMs develop localized, manipulable circuits for meta-cognitive control, enabling test-time steering of computation and reasoning without retraining.} }
Endnote
%0 Conference Paper %T Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking %A Jing Bi %A Luchuan Song %A Dingxin Zhang %A Pinxin Liu %A Guangyu Sun %A Lianggong Bruce Wen %A Weidong Cai %A Chen Chen %A Chenliang Xu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bi26d %I PMLR %P 8063--8075 %U https://proceedings.mlr.press/v306/bi26d.html %V 306 %X Vision-language models (VLMs) now support both direct Instruct and explicit-reasoning Thinking modes, yet practitioners still lack principled ways to decide when reasoning actually improves performance, or how much computation to spend at test time, so we investigate whether VLMs encode meta-cognitive signals for adaptive inference. We derive oracle labels for two properties: (1) reasoning helpfulness, namely whether explicit reasoning improves accuracy, and (2) desired generation length, the minimal token budget needed for a correct answer. Probing final-layer representations in InternVL and Qwen3-VL models, we find Thinking models encode these signals more linearly than Instruct models, suggesting that reasoning-oriented training enhances meta-cognitive structure. Head-wise attribution reveals two circuits: length-control heads in lower layers and reasoning/difficulty heads in higher layers. Causal interventions confirm these roles across benchmarks: scaling length heads controls output length with minimal accuracy loss, while scaling reasoning heads enables a perception-reasoning trade-off, improving accuracy by up to 5.3%. Our results demonstrate that reasoning-tuned VLMs develop localized, manipulable circuits for meta-cognitive control, enabling test-time steering of computation and reasoning without retraining.
APA
Bi, J., Song, L., Zhang, D., Liu, P., Sun, G., Wen, L.B., Cai, W., Chen, C. & Xu, C.. (2026). Does Reasoning Improve Seeing? Understanding When Vision-Language Models Benefit from Thinking. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:8063-8075 Available from https://proceedings.mlr.press/v306/bi26d.html.

Related Material