Psychometric Analysis of MRBench V2

Mayank Sharma
Proceedings of the Impactful and Responsible AI Systems for Education Workshop, PMLR 339:90-100, 2026.

Abstract

Benchmarks for evaluating the pedagogical ability of LLM tutors are increasingly central to educational AI research, yet their psychometric properties are rarely examined formally. We apply a comprehensive measurement validation pipeline to MRBench V2, a benchmark of 200 mathematical tutoring dialogues annotated across eight pedagogical dimensions. Using exploratory and confirmatory factor analyses, graded response modelling, item-level validity diagnostics, measurement invariance testing, and generalizability theory, we find that six of the eight dimensions form a coherent unidimensional scale with excellent structural fit (CFI $= 0.998$, RMSEA $= 0.058$) and strong generalizability ($G_\text{rel} = 0.974$). Two dimensions show psychometric properties inconsistent with their intended role. We further detect measurement non-equivalence across model sizes and highlight open questions regarding gaps in construct validity and predictive validity. Our findings demonstrate the value of psychometric analysis as a standard practice in educational AI evaluation.

Cite this Paper


BibTeX
@InProceedings{pmlr-v339-sharma26a, title = {Psychometric Analysis of MRBench V2}, author = {Sharma, Mayank}, booktitle = {Proceedings of the Impactful and Responsible AI Systems for Education Workshop}, pages = {90--100}, year = {2026}, editor = {Basu Mallick, Debshila and Woodhead, Simon and Wang, Zichao and Ananda, Muktha and Burstein, Jill and Murphy, April}, volume = {339}, series = {Proceedings of Machine Learning Research}, month = {28 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v339/main/assets/sharma26a/sharma26a.pdf}, url = {https://proceedings.mlr.press/v339/sharma26a.html}, abstract = {Benchmarks for evaluating the pedagogical ability of LLM tutors are increasingly central to educational AI research, yet their psychometric properties are rarely examined formally. We apply a comprehensive measurement validation pipeline to MRBench V2, a benchmark of 200 mathematical tutoring dialogues annotated across eight pedagogical dimensions. Using exploratory and confirmatory factor analyses, graded response modelling, item-level validity diagnostics, measurement invariance testing, and generalizability theory, we find that six of the eight dimensions form a coherent unidimensional scale with excellent structural fit (CFI $= 0.998$, RMSEA $= 0.058$) and strong generalizability ($G_\text{rel} = 0.974$). Two dimensions show psychometric properties inconsistent with their intended role. We further detect measurement non-equivalence across model sizes and highlight open questions regarding gaps in construct validity and predictive validity. Our findings demonstrate the value of psychometric analysis as a standard practice in educational AI evaluation.} }
Endnote
%0 Conference Paper %T Psychometric Analysis of MRBench V2 %A Mayank Sharma %B Proceedings of the Impactful and Responsible AI Systems for Education Workshop %C Proceedings of Machine Learning Research %D 2026 %E Debshila Basu Mallick %E Simon Woodhead %E Zichao Wang %E Muktha Ananda %E Jill Burstein %E April Murphy %F pmlr-v339-sharma26a %I PMLR %P 90--100 %U https://proceedings.mlr.press/v339/sharma26a.html %V 339 %X Benchmarks for evaluating the pedagogical ability of LLM tutors are increasingly central to educational AI research, yet their psychometric properties are rarely examined formally. We apply a comprehensive measurement validation pipeline to MRBench V2, a benchmark of 200 mathematical tutoring dialogues annotated across eight pedagogical dimensions. Using exploratory and confirmatory factor analyses, graded response modelling, item-level validity diagnostics, measurement invariance testing, and generalizability theory, we find that six of the eight dimensions form a coherent unidimensional scale with excellent structural fit (CFI $= 0.998$, RMSEA $= 0.058$) and strong generalizability ($G_\text{rel} = 0.974$). Two dimensions show psychometric properties inconsistent with their intended role. We further detect measurement non-equivalence across model sizes and highlight open questions regarding gaps in construct validity and predictive validity. Our findings demonstrate the value of psychometric analysis as a standard practice in educational AI evaluation.
APA
Sharma, M.. (2026). Psychometric Analysis of MRBench V2. Proceedings of the Impactful and Responsible AI Systems for Education Workshop, in Proceedings of Machine Learning Research 339:90-100 Available from https://proceedings.mlr.press/v339/sharma26a.html.

Related Material