STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

Akash Bonagiri, Gerard Janno Anderias, Saee Patil, Angelina Lai, Devang Borkar, Gezheng Kang, Ishant Gandhi, Setareh Rafatirad, Houman Homayoun
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:9039-9057, 2026.

Abstract

Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation. Majority vote discards annotator reliability and item-level ambiguity, often yielding unstable comparisons across annotator subsets. We introduce STABLEVAL, a disagreement-aware evaluation framework that models latent item correctness and annotator-specific confusion patterns to produce posterior expected item credit and calibrated agent-level scores. Unlike label-denoising approaches such as Dawid–Skene, STABLEVAL is explicitly designed for stable and uncertainty-aware system evaluation rather than hard label recovery. We formalize ranking stability as a first-class evaluation objective and analyze how aggregation methods preserve or distort underlying annotator behavior. Across controlled synthetic experiments and multiple real-world human-annotated benchmarks, majority vote exhibits increasing score error and ranking instability under annotator heterogeneity and adversarial noise, while STABLEVAL yields more stable and statistically grounded system rankings. These results demonstrate that modeling disagreement is essential for robust and reproducible AI evaluation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bonagiri26a, title = {{STABLEVAL}: Disagreement-Aware and Stable Evaluation of {AI} Systems}, author = {Bonagiri, Akash and Anderias, Gerard Janno and Patil, Saee and Lai, Angelina and Borkar, Devang and Kang, Gezheng and Gandhi, Ishant and Rafatirad, Setareh and Homayoun, Houman}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {9039--9057}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bonagiri26a/bonagiri26a.pdf}, url = {https://proceedings.mlr.press/v306/bonagiri26a.html}, abstract = {Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation. Majority vote discards annotator reliability and item-level ambiguity, often yielding unstable comparisons across annotator subsets. We introduce STABLEVAL, a disagreement-aware evaluation framework that models latent item correctness and annotator-specific confusion patterns to produce posterior expected item credit and calibrated agent-level scores. Unlike label-denoising approaches such as Dawid–Skene, STABLEVAL is explicitly designed for stable and uncertainty-aware system evaluation rather than hard label recovery. We formalize ranking stability as a first-class evaluation objective and analyze how aggregation methods preserve or distort underlying annotator behavior. Across controlled synthetic experiments and multiple real-world human-annotated benchmarks, majority vote exhibits increasing score error and ranking instability under annotator heterogeneity and adversarial noise, while STABLEVAL yields more stable and statistically grounded system rankings. These results demonstrate that modeling disagreement is essential for robust and reproducible AI evaluation.} }
Endnote
%0 Conference Paper %T STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems %A Akash Bonagiri %A Gerard Janno Anderias %A Saee Patil %A Angelina Lai %A Devang Borkar %A Gezheng Kang %A Ishant Gandhi %A Setareh Rafatirad %A Houman Homayoun %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bonagiri26a %I PMLR %P 9039--9057 %U https://proceedings.mlr.press/v306/bonagiri26a.html %V 306 %X Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation. Majority vote discards annotator reliability and item-level ambiguity, often yielding unstable comparisons across annotator subsets. We introduce STABLEVAL, a disagreement-aware evaluation framework that models latent item correctness and annotator-specific confusion patterns to produce posterior expected item credit and calibrated agent-level scores. Unlike label-denoising approaches such as Dawid–Skene, STABLEVAL is explicitly designed for stable and uncertainty-aware system evaluation rather than hard label recovery. We formalize ranking stability as a first-class evaluation objective and analyze how aggregation methods preserve or distort underlying annotator behavior. Across controlled synthetic experiments and multiple real-world human-annotated benchmarks, majority vote exhibits increasing score error and ranking instability under annotator heterogeneity and adversarial noise, while STABLEVAL yields more stable and statistically grounded system rankings. These results demonstrate that modeling disagreement is essential for robust and reproducible AI evaluation.
APA
Bonagiri, A., Anderias, G.J., Patil, S., Lai, A., Borkar, D., Kang, G., Gandhi, I., Rafatirad, S. & Homayoun, H.. (2026). STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:9039-9057 Available from https://proceedings.mlr.press/v306/bonagiri26a.html.

Related Material