A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

Raghu Arghal, Phoebe Chen, Niall Dalton, Evgenii Kortukov, Calum Mcnamara, Angelos Nalmpantis, Moksh Nirvaan, Gabriele Sarti, Mario Giulianelli
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:3434-3451, 2026.

Abstract

Understanding an agent’s goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models’ internal representations. As a case study, we examine an LLM agent navigating a 2D grid world toward a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues toward immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-arghal26a, title = {A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents}, author = {Arghal, Raghu and Chen, Phoebe and Dalton, Niall and Kortukov, Evgenii and Mcnamara, Calum and Nalmpantis, Angelos and Nirvaan, Moksh and Sarti, Gabriele and Giulianelli, Mario}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {3434--3451}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/arghal26a/arghal26a.pdf}, url = {https://proceedings.mlr.press/v306/arghal26a.html}, abstract = {Understanding an agent’s goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models’ internal representations. As a case study, we examine an LLM agent navigating a 2D grid world toward a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues toward immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.} }
Endnote
%0 Conference Paper %T A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents %A Raghu Arghal %A Phoebe Chen %A Niall Dalton %A Evgenii Kortukov %A Calum Mcnamara %A Angelos Nalmpantis %A Moksh Nirvaan %A Gabriele Sarti %A Mario Giulianelli %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-arghal26a %I PMLR %P 3434--3451 %U https://proceedings.mlr.press/v306/arghal26a.html %V 306 %X Understanding an agent’s goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models’ internal representations. As a case study, we examine an LLM agent navigating a 2D grid world toward a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues toward immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.
APA
Arghal, R., Chen, P., Dalton, N., Kortukov, E., Mcnamara, C., Nalmpantis, A., Nirvaan, M., Sarti, G. & Giulianelli, M.. (2026). A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:3434-3451 Available from https://proceedings.mlr.press/v306/arghal26a.html.

Related Material