CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation

Jitian Zhao, Changho Shin, Tzu-Heng Huang, Satya Sai Srinath Namburi Gnvv, Frederic Sala
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:162114-162150, 2026.

Abstract

LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent estimates of true quality. However, in practice, LLM judges exhibit correlated errors caused by shared latent confounders—such as verbosity, stylistic preferences, or training artifacts—causing standard aggregation rules like majority vote or averaging to provide little gain or even amplify systematic mistakes. To address this, we introduce CARE, a confounder-aware aggregation framework that explicitly models LLM judge scores as arising from both a latent true-quality signal and shared confounding factors. Rather than heuristically re-weighting judges, CARE separates quality from confounders without access to ground-truth labels. We provide theoretical guarantees for identifiability and finite-sample recovery under shared confounders, and we quantify the systematic bias incurred when aggregation models omit confounding latent factors. Across 12 public benchmarks spanning continuous scoring, binary classification, and pairwise preference settings, CARE improves aggregation accuracy, reducing error by up to 26.8%.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-zhao26aq, title = {{CARE}: Confounder-Aware Aggregation for Reliable {LLM} Evaluation}, author = {Zhao, Jitian and Shin, Changho and Huang, Tzu-Heng and Namburi Gnvv, Satya Sai Srinath and Sala, Frederic}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {162114--162150}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/zhao26aq/zhao26aq.pdf}, url = {https://proceedings.mlr.press/v306/zhao26aq.html}, abstract = {LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent estimates of true quality. However, in practice, LLM judges exhibit correlated errors caused by shared latent confounders—such as verbosity, stylistic preferences, or training artifacts—causing standard aggregation rules like majority vote or averaging to provide little gain or even amplify systematic mistakes. To address this, we introduce CARE, a confounder-aware aggregation framework that explicitly models LLM judge scores as arising from both a latent true-quality signal and shared confounding factors. Rather than heuristically re-weighting judges, CARE separates quality from confounders without access to ground-truth labels. We provide theoretical guarantees for identifiability and finite-sample recovery under shared confounders, and we quantify the systematic bias incurred when aggregation models omit confounding latent factors. Across 12 public benchmarks spanning continuous scoring, binary classification, and pairwise preference settings, CARE improves aggregation accuracy, reducing error by up to 26.8%.} }
Endnote
%0 Conference Paper %T CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation %A Jitian Zhao %A Changho Shin %A Tzu-Heng Huang %A Satya Sai Srinath Namburi Gnvv %A Frederic Sala %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-zhao26aq %I PMLR %P 162114--162150 %U https://proceedings.mlr.press/v306/zhao26aq.html %V 306 %X LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges provide independent estimates of true quality. However, in practice, LLM judges exhibit correlated errors caused by shared latent confounders—such as verbosity, stylistic preferences, or training artifacts—causing standard aggregation rules like majority vote or averaging to provide little gain or even amplify systematic mistakes. To address this, we introduce CARE, a confounder-aware aggregation framework that explicitly models LLM judge scores as arising from both a latent true-quality signal and shared confounding factors. Rather than heuristically re-weighting judges, CARE separates quality from confounders without access to ground-truth labels. We provide theoretical guarantees for identifiability and finite-sample recovery under shared confounders, and we quantify the systematic bias incurred when aggregation models omit confounding latent factors. Across 12 public benchmarks spanning continuous scoring, binary classification, and pairwise preference settings, CARE improves aggregation accuracy, reducing error by up to 26.8%.
APA
Zhao, J., Shin, C., Huang, T., Namburi Gnvv, S.S.S. & Sala, F.. (2026). CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:162114-162150 Available from https://proceedings.mlr.press/v306/zhao26aq.html.

Related Material