Quantifying Biases in LLM-as-a-Judge Evaluations

Magda Dubois, Harry Coppock, Mario Giulianelli, Ole Kristian Jorgensen, Timo Flesch, Lennart Luettgau, Cozmin Ududec
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27120-27135, 2026.

Abstract

The evaluation of large language models (LLMs) is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they are not free from biases (e.g., favouring longer outputs or generations from their own model family). Here we propose a statistical framework based on Bayesian generalised linear models (GLMs) that enables researchers to address their primary research questions (e.g., LLM capability or risk assessment), while simultaneously identifying, quantifying and mitigating various biases in their autograders. Our approach can be applied to various evaluation formats (e.g., absolute scores or pairwise preferences) and augments traditional metrics (e.g., inter-rater agreement) by providing precise uncertainty estimates and clarifying sources of disagreement between graders. This framework also enables efficient counterfactual simulations without costly re-evaluation (e.g., assessing agreement after removing systematic biases). We demonstrate these capabilities through simulated examples, with all methods available in an open-source software package. Overall, we introduce a novel framework for autograder evaluation which allows researchers to detect, quantify and correct for various biases in a systematic way.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-dubois26a, title = {Quantifying Biases in {LLM}-as-a-Judge Evaluations}, author = {Dubois, Magda and Coppock, Harry and Giulianelli, Mario and Jorgensen, Ole Kristian and Flesch, Timo and Luettgau, Lennart and Ududec, Cozmin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27120--27135}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/dubois26a/dubois26a.pdf}, url = {https://proceedings.mlr.press/v306/dubois26a.html}, abstract = {The evaluation of large language models (LLMs) is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they are not free from biases (e.g., favouring longer outputs or generations from their own model family). Here we propose a statistical framework based on Bayesian generalised linear models (GLMs) that enables researchers to address their primary research questions (e.g., LLM capability or risk assessment), while simultaneously identifying, quantifying and mitigating various biases in their autograders. Our approach can be applied to various evaluation formats (e.g., absolute scores or pairwise preferences) and augments traditional metrics (e.g., inter-rater agreement) by providing precise uncertainty estimates and clarifying sources of disagreement between graders. This framework also enables efficient counterfactual simulations without costly re-evaluation (e.g., assessing agreement after removing systematic biases). We demonstrate these capabilities through simulated examples, with all methods available in an open-source software package. Overall, we introduce a novel framework for autograder evaluation which allows researchers to detect, quantify and correct for various biases in a systematic way.} }
Endnote
%0 Conference Paper %T Quantifying Biases in LLM-as-a-Judge Evaluations %A Magda Dubois %A Harry Coppock %A Mario Giulianelli %A Ole Kristian Jorgensen %A Timo Flesch %A Lennart Luettgau %A Cozmin Ududec %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-dubois26a %I PMLR %P 27120--27135 %U https://proceedings.mlr.press/v306/dubois26a.html %V 306 %X The evaluation of large language models (LLMs) is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they are not free from biases (e.g., favouring longer outputs or generations from their own model family). Here we propose a statistical framework based on Bayesian generalised linear models (GLMs) that enables researchers to address their primary research questions (e.g., LLM capability or risk assessment), while simultaneously identifying, quantifying and mitigating various biases in their autograders. Our approach can be applied to various evaluation formats (e.g., absolute scores or pairwise preferences) and augments traditional metrics (e.g., inter-rater agreement) by providing precise uncertainty estimates and clarifying sources of disagreement between graders. This framework also enables efficient counterfactual simulations without costly re-evaluation (e.g., assessing agreement after removing systematic biases). We demonstrate these capabilities through simulated examples, with all methods available in an open-source software package. Overall, we introduce a novel framework for autograder evaluation which allows researchers to detect, quantify and correct for various biases in a systematic way.
APA
Dubois, M., Coppock, H., Giulianelli, M., Jorgensen, O.K., Flesch, T., Luettgau, L. & Ududec, C.. (2026). Quantifying Biases in LLM-as-a-Judge Evaluations. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27120-27135 Available from https://proceedings.mlr.press/v306/dubois26a.html.

Related Material