[edit]
Beyond Aggregate Accuracy: Continuous Evaluation of Residual Failures in Image-Required Educational Mathematics
Proceedings of the Impactful and Responsible AI Systems for Education Workshop, PMLR 339:170-179, 2026.
Abstract
Rapid model progress can make educational benchmark results stale, but responsible evaluation requires more than refreshing aggregate scores. We revisit a benchmark of 376 curriculum-authentic, image-required middle-school mathematics items using GPT-5.5 and compare its behavior with representative prior models. In the \textit{with-images} condition, accuracy increases from about 88% for GPT-5.4 to about 93% for GPT-5.5, while the strict never-solved set decreases from 31 items to 17. In contrast, \textit{without-images} performance remains near zero, confirming that the benchmark remains genuinely image-required. We then audit this residual set and find that its composition clarifies what the remaining failures mean: they include inaccessible-image cases, insufficient-context items, answer-format mismatches, and benchmark-quality issues rather than only generic visual-reasoning failures. We treat representation-sensitive probes as future work, using the audit to identify which residual items are appropriate probe candidates. We argue that continuous evaluation of educational AI systems should pair aggregate performance updates with residual item audits that distinguish model limitations from dataset, scoring, accessibility, and task-specification issues.