[edit]
Co-TA: Streamlining Automated Grading with Generative Rubrics and Distribution Simulation
Proceedings of the Impactful and Responsible AI Systems for Education Workshop, PMLR 339:146-159, 2026.
Abstract
Large Language Models (LLMs) offer great potential to reduce Teaching Assistant workloads through automated grading. However, their rapid integration shifts the primary challenge from technical scalability to responsible governance, as LLMs frequently diverge from expert human graders on complex rubrics. Existing Human-in-the-Loop (HITL) systems focus on micro-level calibration for individual submissions, leaving instructors blind to how rubrics alter class-wide grade distributions. To address this gap, we introduce \textbf{Co-TA}, a deployable web application that streamlines the grading workflow while shifting the paradigm to rubric co-design and fairness auditing. Featuring batch uploads and uncertainty flagging, Co-TA allows instructors to generate rubric variations and immediately simulate their impact against synthetic student personas. By visualizing resulting grade distributions prior to live grading, Co-TA serves as a critical fairness auditing mechanism, ensuring AI assessment remains transparent, verifiable, and firmly under human pedagogical control.