PRISM: Perception Reasoning Interleaved for Sequential Decision Making.

Mohamed Salim Aissi, Clémence Grislain, Clément Romac, Laure Soulier, Mohamed Chetouani, Olivier Sigaud, Nicolas Thome
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:1451-1474, 2026.

Abstract

Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception–reasoning–decision gap in standalone Vision–Language Models (VLMs), which often overlook task-critical information. In this paper, we introduce PRISM, a framework that tightly couples perception (VLM) and decision (LLM) through a dynamic question–answer (DQA) pipeline. Instead of passively accepting the VLM’s description, the LLM critiques it, probes the VLM with goal-oriented questions, and synthesizes a compact image description. This closed-loop interaction yields a sharp, task-driven understanding of the scene. We evaluate PRISM on the ALFWorld and Room-to-Room (R2R) benchmarks. We show that: (1) PRISM significantly outperforms state-of-the-art image-based models, (2) our Interactive goal-oriented perception pipeline yields systematic and substantial gains, and (3) PRISM is fully automatic, eliminating the need for handcrafted questions or answers.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-aissi26a, title = {{PRISM}: Perception Reasoning Interleaved for Sequential Decision Making.}, author = {Aissi, Mohamed Salim and Grislain, Cl\'{e}mence and Romac, Cl\'{e}ment and Soulier, Laure and Chetouani, Mohamed and Sigaud, Olivier and Thome, Nicolas}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {1451--1474}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/aissi26a/aissi26a.pdf}, url = {https://proceedings.mlr.press/v306/aissi26a.html}, abstract = {Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception–reasoning–decision gap in standalone Vision–Language Models (VLMs), which often overlook task-critical information. In this paper, we introduce PRISM, a framework that tightly couples perception (VLM) and decision (LLM) through a dynamic question–answer (DQA) pipeline. Instead of passively accepting the VLM’s description, the LLM critiques it, probes the VLM with goal-oriented questions, and synthesizes a compact image description. This closed-loop interaction yields a sharp, task-driven understanding of the scene. We evaluate PRISM on the ALFWorld and Room-to-Room (R2R) benchmarks. We show that: (1) PRISM significantly outperforms state-of-the-art image-based models, (2) our Interactive goal-oriented perception pipeline yields systematic and substantial gains, and (3) PRISM is fully automatic, eliminating the need for handcrafted questions or answers.} }
Endnote
%0 Conference Paper %T PRISM: Perception Reasoning Interleaved for Sequential Decision Making. %A Mohamed Salim Aissi %A Clémence Grislain %A Clément Romac %A Laure Soulier %A Mohamed Chetouani %A Olivier Sigaud %A Nicolas Thome %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-aissi26a %I PMLR %P 1451--1474 %U https://proceedings.mlr.press/v306/aissi26a.html %V 306 %X Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception–reasoning–decision gap in standalone Vision–Language Models (VLMs), which often overlook task-critical information. In this paper, we introduce PRISM, a framework that tightly couples perception (VLM) and decision (LLM) through a dynamic question–answer (DQA) pipeline. Instead of passively accepting the VLM’s description, the LLM critiques it, probes the VLM with goal-oriented questions, and synthesizes a compact image description. This closed-loop interaction yields a sharp, task-driven understanding of the scene. We evaluate PRISM on the ALFWorld and Room-to-Room (R2R) benchmarks. We show that: (1) PRISM significantly outperforms state-of-the-art image-based models, (2) our Interactive goal-oriented perception pipeline yields systematic and substantial gains, and (3) PRISM is fully automatic, eliminating the need for handcrafted questions or answers.
APA
Aissi, M.S., Grislain, C., Romac, C., Soulier, L., Chetouani, M., Sigaud, O. & Thome, N.. (2026). PRISM: Perception Reasoning Interleaved for Sequential Decision Making.. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:1451-1474 Available from https://proceedings.mlr.press/v306/aissi26a.html.

Related Material