[edit]
Evaluating the Landscape of Automated Scoring with GenAI: A Systematic Review
Proceedings of the Impactful and Responsible AI Systems for Education Workshop, PMLR 339:140-145, 2026.
Abstract
As Generative Artificial Intelligence (GenAI) capabilities rapidly expand in educational settings, assessing the validity, reliability, and fairness of these tools is critical before widespread classroom deployment. While there has been a large amount of research, it is not yet clear what methods are being used to evaluate these foundational concepts and what the findings are. We present a work-in-progress systematic literature review focusing exclusively on automated item scoring. Based on a final coded corpus of 107 empirical studies published since January 2023, our preliminary analysis reveals that while prompt-based Large Language Models (LLMs) are achieving parity with traditional transformer-based methods, significant gaps remain regarding fairness evaluation and the representation of diverse student populations. Quadratic weighted kappa (QWK) is the dominant agreement metric, used in 41.8% of studies, and fairness is acknowledged by 55% of studies but empirically evaluated in only 11.2%. This paper offers evidence-based guidance for responsible classroom integration of AI.