QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye, Edward Zhang, Ruslans Aleksejevs, Todor Antić, Polina Baron, Sujeet Bhalerao, Shubhrajit Bhattacharya, Zachary Burton, John Byrne, Hyungjun Choi, Nujhat Ahmed Disha, Koppány István Encz, Yuchen Fang, Robert Joseph George, Ebrahim Ghorbani, Alan Goldfarb, Jing Guo, Meghal Gupta, Stefano Huber, Annika Kanckos, Minjung Kang, Hyun Jong Kim, Dino Lorenzini, Levi Lorenzo, Tianyi Mao, Giovanni Marzenta, Ariane M. Masuda, Lukas Mauth, Ana Mickovic, Andrés Miniguano-Trujillo, Antoine Moulin, Wenqi Ni, Tomos Parry, Kevin Ren, Hossein Roodbarani, Mathieu Rundström, Manjil Saikia, Detchat Samart, Rebecca Steiner, Connor Stewart, Dhara Thakkar, Jeffrey Tse, Vasiliki Velona, Yunhai Xiang, Sibel Yalçın, Jun Yan, Ji Zeng, Arman Cohan, Quanquan C. Liu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:36056-36179, 2026.

Abstract

As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic evaluation Alignment Gap when applied to upper-undergraduate to early graduate level mathematics. To quantify this, we introduce QEDBench, the first benchmark to systematically measure alignment with human experts on undergraduate-level math proofs by contrasting course-specific rubrics against expert common knowledge criteria. By deploying a dual-evaluation matrix ($7$ judges $\times$ $5$ solvers) against 1,000+ hours of human evaluation, we reveal that certain frontier evaluators like Claude 4.5 Opus exhibit significant positive bias (up to $+0.28$ mean score inflation), effectively "hallucinating rigor" in flawed proofs. Furthermore, we uncover a critical reasoning disparity: while Gemini 3.0 Pro achieves state-of-the-art performance (0.91 raw score), specialized reasoning models like o3-deep-research collapse in discrete domains, dropping to 42.1% accuracy in Graph Theory. We release QEDBench as a public benchmark for evaluating and improving AI judges.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gonzalez26a, title = {{QEDB}ench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs}, author = {Gonzalez, Santiago and Bavandpour, Alireza Amiri and Ye, Peter and Zhang, Edward and Aleksejevs, Ruslans and Anti\'{c}, Todor and Baron, Polina and Bhalerao, Sujeet and Bhattacharya, Shubhrajit and Burton, Zachary and Byrne, John and Choi, Hyungjun and Disha, Nujhat Ahmed and Encz, Kopp\'{a}ny Istv\'{a}n and Fang, Yuchen and George, Robert Joseph and Ghorbani, Ebrahim and Goldfarb, Alan and Guo, Jing and Gupta, Meghal and Huber, Stefano and Kanckos, Annika and Kang, Minjung and Kim, Hyun Jong and Lorenzini, Dino and Lorenzo, Levi and Mao, Tianyi and Marzenta, Giovanni and Masuda, Ariane M. and Mauth, Lukas and Mickovic, Ana and Miniguano-Trujillo, Andr\'{e}s and Moulin, Antoine and Ni, Wenqi and Parry, Tomos and Ren, Kevin and Roodbarani, Hossein and Rundstr\"{o}m, Mathieu and Saikia, Manjil and Samart, Detchat and Steiner, Rebecca and Stewart, Connor and Thakkar, Dhara and Tse, Jeffrey and Velona, Vasiliki and Xiang, Yunhai and Yal\c{c}{\i}n, Sibel and Yan, Jun and Zeng, Ji and Cohan, Arman and Liu, Quanquan C.}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {36056--36179}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gonzalez26a/gonzalez26a.pdf}, url = {https://proceedings.mlr.press/v306/gonzalez26a.html}, abstract = {As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic evaluation Alignment Gap when applied to upper-undergraduate to early graduate level mathematics. To quantify this, we introduce QEDBench, the first benchmark to systematically measure alignment with human experts on undergraduate-level math proofs by contrasting course-specific rubrics against expert common knowledge criteria. By deploying a dual-evaluation matrix ($7$ judges $\times$ $5$ solvers) against 1,000+ hours of human evaluation, we reveal that certain frontier evaluators like Claude 4.5 Opus exhibit significant positive bias (up to $+0.28$ mean score inflation), effectively "hallucinating rigor" in flawed proofs. Furthermore, we uncover a critical reasoning disparity: while Gemini 3.0 Pro achieves state-of-the-art performance (0.91 raw score), specialized reasoning models like o3-deep-research collapse in discrete domains, dropping to 42.1% accuracy in Graph Theory. We release QEDBench as a public benchmark for evaluating and improving AI judges.} }
Endnote
%0 Conference Paper %T QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs %A Santiago Gonzalez %A Alireza Amiri Bavandpour %A Peter Ye %A Edward Zhang %A Ruslans Aleksejevs %A Todor Antić %A Polina Baron %A Sujeet Bhalerao %A Shubhrajit Bhattacharya %A Zachary Burton %A John Byrne %A Hyungjun Choi %A Nujhat Ahmed Disha %A Koppány István Encz %A Yuchen Fang %A Robert Joseph George %A Ebrahim Ghorbani %A Alan Goldfarb %A Jing Guo %A Meghal Gupta %A Stefano Huber %A Annika Kanckos %A Minjung Kang %A Hyun Jong Kim %A Dino Lorenzini %A Levi Lorenzo %A Tianyi Mao %A Giovanni Marzenta %A Ariane M. Masuda %A Lukas Mauth %A Ana Mickovic %A Andrés Miniguano-Trujillo %A Antoine Moulin %A Wenqi Ni %A Tomos Parry %A Kevin Ren %A Hossein Roodbarani %A Mathieu Rundström %A Manjil Saikia %A Detchat Samart %A Rebecca Steiner %A Connor Stewart %A Dhara Thakkar %A Jeffrey Tse %A Vasiliki Velona %A Yunhai Xiang %A Sibel Yalçın %A Jun Yan %A Ji Zeng %A Arman Cohan %A Quanquan C. Liu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gonzalez26a %I PMLR %P 36056--36179 %U https://proceedings.mlr.press/v306/gonzalez26a.html %V 306 %X As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic evaluation Alignment Gap when applied to upper-undergraduate to early graduate level mathematics. To quantify this, we introduce QEDBench, the first benchmark to systematically measure alignment with human experts on undergraduate-level math proofs by contrasting course-specific rubrics against expert common knowledge criteria. By deploying a dual-evaluation matrix ($7$ judges $\times$ $5$ solvers) against 1,000+ hours of human evaluation, we reveal that certain frontier evaluators like Claude 4.5 Opus exhibit significant positive bias (up to $+0.28$ mean score inflation), effectively "hallucinating rigor" in flawed proofs. Furthermore, we uncover a critical reasoning disparity: while Gemini 3.0 Pro achieves state-of-the-art performance (0.91 raw score), specialized reasoning models like o3-deep-research collapse in discrete domains, dropping to 42.1% accuracy in Graph Theory. We release QEDBench as a public benchmark for evaluating and improving AI judges.
APA
Gonzalez, S., Bavandpour, A.A., Ye, P., Zhang, E., Aleksejevs, R., Antić, T., Baron, P., Bhalerao, S., Bhattacharya, S., Burton, Z., Byrne, J., Choi, H., Disha, N.A., Encz, K.I., Fang, Y., George, R.J., Ghorbani, E., Goldfarb, A., Guo, J., Gupta, M., Huber, S., Kanckos, A., Kang, M., Kim, H.J., Lorenzini, D., Lorenzo, L., Mao, T., Marzenta, G., Masuda, A.M., Mauth, L., Mickovic, A., Miniguano-Trujillo, A., Moulin, A., Ni, W., Parry, T., Ren, K., Roodbarani, H., Rundström, M., Saikia, M., Samart, D., Steiner, R., Stewart, C., Thakkar, D., Tse, J., Velona, V., Xiang, Y., Yalçın, S., Yan, J., Zeng, J., Cohan, A. & Liu, Q.C.. (2026). QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:36056-36179 Available from https://proceedings.mlr.press/v306/gonzalez26a.html.

Related Material