Should You Use Your Large Language Model to Explore or Exploit?

Keegan Harris, Aleksandrs Slivkins
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2008-2058, 2026.

Abstract

We evaluate the ability of the current generation of large language models ({LLMs}) to help a decision-making agent facing an exploration-exploitation tradeoff. While previous work has largely study the ability of {LLMs} to solve combined exploration-exploitation tasks, we take a more systematic approach and use {LLMs} to explore and exploit in silos in various (contextual) bandit tasks. We find that reasoning models show the most promise for solving exploitation tasks, although they are still too expensive or too slow to be used in many practical settings. Motivated by this, we study tool use and in-context summarization using non-reasoning models. We find that these mitigations may be used to substantially improve performance on medium-difficulty tasks, however even then, all {LLMs} we study perform worse than a simple linear regression, even in non-linear settings. On the other hand, we find that {LLMs} do help at exploring large action spaces with inherent semantics, by suggesting suitable candidates to explore.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-harris26a, title = {Should You Use Your Large Language Model to Explore or Exploit?}, author = {Harris, Keegan and Slivkins, Aleksandrs}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {2008--2058}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/harris26a/harris26a.pdf}, url = {https://proceedings.mlr.press/v337/harris26a.html}, abstract = {We evaluate the ability of the current generation of large language models ({LLMs}) to help a decision-making agent facing an exploration-exploitation tradeoff. While previous work has largely study the ability of {LLMs} to solve combined exploration-exploitation tasks, we take a more systematic approach and use {LLMs} to explore and exploit in silos in various (contextual) bandit tasks. We find that reasoning models show the most promise for solving exploitation tasks, although they are still too expensive or too slow to be used in many practical settings. Motivated by this, we study tool use and in-context summarization using non-reasoning models. We find that these mitigations may be used to substantially improve performance on medium-difficulty tasks, however even then, all {LLMs} we study perform worse than a simple linear regression, even in non-linear settings. On the other hand, we find that {LLMs} do help at exploring large action spaces with inherent semantics, by suggesting suitable candidates to explore.} }
Endnote
%0 Conference Paper %T Should You Use Your Large Language Model to Explore or Exploit? %A Keegan Harris %A Aleksandrs Slivkins %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-harris26a %I PMLR %P 2008--2058 %U https://proceedings.mlr.press/v337/harris26a.html %V 337 %X We evaluate the ability of the current generation of large language models ({LLMs}) to help a decision-making agent facing an exploration-exploitation tradeoff. While previous work has largely study the ability of {LLMs} to solve combined exploration-exploitation tasks, we take a more systematic approach and use {LLMs} to explore and exploit in silos in various (contextual) bandit tasks. We find that reasoning models show the most promise for solving exploitation tasks, although they are still too expensive or too slow to be used in many practical settings. Motivated by this, we study tool use and in-context summarization using non-reasoning models. We find that these mitigations may be used to substantially improve performance on medium-difficulty tasks, however even then, all {LLMs} we study perform worse than a simple linear regression, even in non-linear settings. On the other hand, we find that {LLMs} do help at exploring large action spaces with inherent semantics, by suggesting suitable candidates to explore.
APA
Harris, K. & Slivkins, A.. (2026). Should You Use Your Large Language Model to Explore or Exploit?. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:2008-2058 Available from https://proceedings.mlr.press/v337/harris26a.html.

Related Material