Intrinsic Credit Assignment for Long Horizon Interaction

Ilze Amanda Auzina, Joschka Strüber, Sergio Hernández-Gutiérrez, Shashwat Goel, Ameya Prabhu, Matthias Bethge
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:4328-4384, 2026.

Abstract

How can we train agents to navigate uncertainty over long horizons? In this work, we propose $\Delta$Belief-RL, which leverages a language model’s own intrinsic beliefs to reward intermediate progress. Our method utilizes the change in the probability an agent assigns to the target solution for credit assignment. By training on synthetic interaction data, $\Delta$Belief-RL teaches information-seeking capabilities that consistently outperform purely outcome-based rewards for RL, with improvements generalizing to out-of-distribution applications ranging from customer service to personalization. Notably, the performance continues to improve as we scale test-time interactions beyond the training horizon, with interaction-efficiency increasing even on Pass@k metrics. Overall, our work introduces a scalable training strategy for navigating uncertainty over a long-horizon, by enabling credit assignment to intermediate actions via intrinsic $\Delta$Belief rewards.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-auzina26a, title = {Intrinsic Credit Assignment for Long Horizon Interaction}, author = {Auzina, Ilze Amanda and Str\"{u}ber, Joschka and Hern\'{a}ndez-Guti\'{e}rrez, Sergio and Goel, Shashwat and Prabhu, Ameya and Bethge, Matthias}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {4328--4384}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/auzina26a/auzina26a.pdf}, url = {https://proceedings.mlr.press/v306/auzina26a.html}, abstract = {How can we train agents to navigate uncertainty over long horizons? In this work, we propose $\Delta$Belief-RL, which leverages a language model’s own intrinsic beliefs to reward intermediate progress. Our method utilizes the change in the probability an agent assigns to the target solution for credit assignment. By training on synthetic interaction data, $\Delta$Belief-RL teaches information-seeking capabilities that consistently outperform purely outcome-based rewards for RL, with improvements generalizing to out-of-distribution applications ranging from customer service to personalization. Notably, the performance continues to improve as we scale test-time interactions beyond the training horizon, with interaction-efficiency increasing even on Pass@k metrics. Overall, our work introduces a scalable training strategy for navigating uncertainty over a long-horizon, by enabling credit assignment to intermediate actions via intrinsic $\Delta$Belief rewards.} }
Endnote
%0 Conference Paper %T Intrinsic Credit Assignment for Long Horizon Interaction %A Ilze Amanda Auzina %A Joschka Strüber %A Sergio Hernández-Gutiérrez %A Shashwat Goel %A Ameya Prabhu %A Matthias Bethge %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-auzina26a %I PMLR %P 4328--4384 %U https://proceedings.mlr.press/v306/auzina26a.html %V 306 %X How can we train agents to navigate uncertainty over long horizons? In this work, we propose $\Delta$Belief-RL, which leverages a language model’s own intrinsic beliefs to reward intermediate progress. Our method utilizes the change in the probability an agent assigns to the target solution for credit assignment. By training on synthetic interaction data, $\Delta$Belief-RL teaches information-seeking capabilities that consistently outperform purely outcome-based rewards for RL, with improvements generalizing to out-of-distribution applications ranging from customer service to personalization. Notably, the performance continues to improve as we scale test-time interactions beyond the training horizon, with interaction-efficiency increasing even on Pass@k metrics. Overall, our work introduces a scalable training strategy for navigating uncertainty over a long-horizon, by enabling credit assignment to intermediate actions via intrinsic $\Delta$Belief rewards.
APA
Auzina, I.A., Strüber, J., Hernández-Gutiérrez, S., Goel, S., Prabhu, A. & Bethge, M.. (2026). Intrinsic Credit Assignment for Long Horizon Interaction. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:4328-4384 Available from https://proceedings.mlr.press/v306/auzina26a.html.

Related Material