Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

Wei Deng, Xianlin Zhang, Mengshi Qi
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:24407-24417, 2026.

Abstract

Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is difficult for real-world applications. Moreover, reinforcement learning methods rely on sparse rewards, limiting their effectiveness for complex reasoning tasks. Inspired by pigeons’ building and exploiting cognitive maps for navigation, we propose a novel agentic pipeline for spatial reasoning. First, we introduce a new dynamic cognitive map parameterizing scene layout as object positions and orientations, serving as persistent memory for new observations. Second, we propose a novel Spatial Assertion Codes (SAC), Python expressions programmatically describing spatial relationships. By collaborating with the dynamic cognitive map, SAC enables verification of intermediate reasoning steps, providing dense reward signals. We optimize the model via supervised and reinforcement finetuning. Experiments on the MindCube benchmark demonstrate state-of-the-art performance with 80.5% overall accuracy, outperforming the best current method by 29.5 accuracy points (a relative improvement of 53.2%) on the challenging Rotation subset. Our code and data are open-sourced at https://github.com/dw-dengwei/active-spatial-reasoning.git.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-deng26x, title = {Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models}, author = {Deng, Wei and Zhang, Xianlin and Qi, Mengshi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {24407--24417}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/deng26x/deng26x.pdf}, url = {https://proceedings.mlr.press/v306/deng26x.html}, abstract = {Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is difficult for real-world applications. Moreover, reinforcement learning methods rely on sparse rewards, limiting their effectiveness for complex reasoning tasks. Inspired by pigeons’ building and exploiting cognitive maps for navigation, we propose a novel agentic pipeline for spatial reasoning. First, we introduce a new dynamic cognitive map parameterizing scene layout as object positions and orientations, serving as persistent memory for new observations. Second, we propose a novel Spatial Assertion Codes (SAC), Python expressions programmatically describing spatial relationships. By collaborating with the dynamic cognitive map, SAC enables verification of intermediate reasoning steps, providing dense reward signals. We optimize the model via supervised and reinforcement finetuning. Experiments on the MindCube benchmark demonstrate state-of-the-art performance with 80.5% overall accuracy, outperforming the best current method by 29.5 accuracy points (a relative improvement of 53.2%) on the challenging Rotation subset. Our code and data are open-sourced at https://github.com/dw-dengwei/active-spatial-reasoning.git.} }
Endnote
%0 Conference Paper %T Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models %A Wei Deng %A Xianlin Zhang %A Mengshi Qi %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-deng26x %I PMLR %P 24407--24417 %U https://proceedings.mlr.press/v306/deng26x.html %V 306 %X Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is difficult for real-world applications. Moreover, reinforcement learning methods rely on sparse rewards, limiting their effectiveness for complex reasoning tasks. Inspired by pigeons’ building and exploiting cognitive maps for navigation, we propose a novel agentic pipeline for spatial reasoning. First, we introduce a new dynamic cognitive map parameterizing scene layout as object positions and orientations, serving as persistent memory for new observations. Second, we propose a novel Spatial Assertion Codes (SAC), Python expressions programmatically describing spatial relationships. By collaborating with the dynamic cognitive map, SAC enables verification of intermediate reasoning steps, providing dense reward signals. We optimize the model via supervised and reinforcement finetuning. Experiments on the MindCube benchmark demonstrate state-of-the-art performance with 80.5% overall accuracy, outperforming the best current method by 29.5 accuracy points (a relative improvement of 53.2%) on the challenging Rotation subset. Our code and data are open-sourced at https://github.com/dw-dengwei/active-spatial-reasoning.git.
APA
Deng, W., Zhang, X. & Qi, M.. (2026). Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:24407-24417 Available from https://proceedings.mlr.press/v306/deng26x.html.

Related Material