Evaluation of Large Language Models via Coupled Token Generation

Nina L. Corvelo Benz, Stratis Tsirtsis, Eleni Straitouri, Ivi Chatzi, Ander Artola Velasco, Suhas Thejaswi, Manuel Gomez Rodriguez
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2215-2223, 2026.

Abstract

State-of-the-art large language models rely on randomization to respond to a prompt. Consequently, a model may respond differently to the same prompt if asked multiple times. In this work, we argue that the evaluation and ranking of large language models should control for this randomization. Our starting point is the development of a causal model for coupled autoregressive generation, which allows different large language models to sample responses with the same source of randomness. Building upon our causal model, we first show that, on evaluations based on benchmark datasets, coupled autoregressive generation leads to the same conclusions as vanilla autoregressive generation but using provably fewer samples. However, we further show that, on evaluations based on pairwise comparisons, the two approaches can surprisingly lead to different rankings when comparing more than two models. To complement our theoretical results, we conduct experiments with several models from the $\texttt{Llama}$, $\texttt{Mistral}$ and $\texttt{Qwen}$ families. We find that, across multiple benchmark datasets, coupled autoregressive generation requires up to $75$% fewer samples to reach the same conclusions as vanilla autoregressive generation. Further, we find that the win-rates derived from pairwise comparisons by a strong large language model to prompts from the LMSYS Chatbot Arena platform differ under coupled and vanilla autoregressive generation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-benz26a, title = { Evaluation of Large Language Models via Coupled Token Generation }, author = {Benz, Nina L. Corvelo and Tsirtsis, Stratis and Straitouri, Eleni and Chatzi, Ivi and Velasco, Ander Artola and Thejaswi, Suhas and Rodriguez, Manuel Gomez}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2215--2223}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/benz26a/benz26a.pdf}, url = {https://proceedings.mlr.press/v300/benz26a.html}, abstract = { State-of-the-art large language models rely on randomization to respond to a prompt. Consequently, a model may respond differently to the same prompt if asked multiple times. In this work, we argue that the evaluation and ranking of large language models should control for this randomization. Our starting point is the development of a causal model for coupled autoregressive generation, which allows different large language models to sample responses with the same source of randomness. Building upon our causal model, we first show that, on evaluations based on benchmark datasets, coupled autoregressive generation leads to the same conclusions as vanilla autoregressive generation but using provably fewer samples. However, we further show that, on evaluations based on pairwise comparisons, the two approaches can surprisingly lead to different rankings when comparing more than two models. To complement our theoretical results, we conduct experiments with several models from the $\texttt{Llama}$, $\texttt{Mistral}$ and $\texttt{Qwen}$ families. We find that, across multiple benchmark datasets, coupled autoregressive generation requires up to $75$% fewer samples to reach the same conclusions as vanilla autoregressive generation. Further, we find that the win-rates derived from pairwise comparisons by a strong large language model to prompts from the LMSYS Chatbot Arena platform differ under coupled and vanilla autoregressive generation. } }
Endnote
%0 Conference Paper %T Evaluation of Large Language Models via Coupled Token Generation %A Nina L. Corvelo Benz %A Stratis Tsirtsis %A Eleni Straitouri %A Ivi Chatzi %A Ander Artola Velasco %A Suhas Thejaswi %A Manuel Gomez Rodriguez %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-benz26a %I PMLR %P 2215--2223 %U https://proceedings.mlr.press/v300/benz26a.html %V 300 %X State-of-the-art large language models rely on randomization to respond to a prompt. Consequently, a model may respond differently to the same prompt if asked multiple times. In this work, we argue that the evaluation and ranking of large language models should control for this randomization. Our starting point is the development of a causal model for coupled autoregressive generation, which allows different large language models to sample responses with the same source of randomness. Building upon our causal model, we first show that, on evaluations based on benchmark datasets, coupled autoregressive generation leads to the same conclusions as vanilla autoregressive generation but using provably fewer samples. However, we further show that, on evaluations based on pairwise comparisons, the two approaches can surprisingly lead to different rankings when comparing more than two models. To complement our theoretical results, we conduct experiments with several models from the $\texttt{Llama}$, $\texttt{Mistral}$ and $\texttt{Qwen}$ families. We find that, across multiple benchmark datasets, coupled autoregressive generation requires up to $75$% fewer samples to reach the same conclusions as vanilla autoregressive generation. Further, we find that the win-rates derived from pairwise comparisons by a strong large language model to prompts from the LMSYS Chatbot Arena platform differ under coupled and vanilla autoregressive generation.
APA
Benz, N.L.C., Tsirtsis, S., Straitouri, E., Chatzi, I., Velasco, A.A., Thejaswi, S. & Rodriguez, M.G.. (2026). Evaluation of Large Language Models via Coupled Token Generation . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2215-2223 Available from https://proceedings.mlr.press/v300/benz26a.html.

Related Material