Co-TA: Streamlining Automated Grading with Generative Rubrics and Distribution Simulation

Helen Jin, Albert Chen
Proceedings of the Impactful and Responsible AI Systems for Education Workshop, PMLR 339:146-159, 2026.

Abstract

Large Language Models (LLMs) offer great potential to reduce Teaching Assistant workloads through automated grading. However, their rapid integration shifts the primary challenge from technical scalability to responsible governance, as LLMs frequently diverge from expert human graders on complex rubrics. Existing Human-in-the-Loop (HITL) systems focus on micro-level calibration for individual submissions, leaving instructors blind to how rubrics alter class-wide grade distributions. To address this gap, we introduce \textbf{Co-TA}, a deployable web application that streamlines the grading workflow while shifting the paradigm to rubric co-design and fairness auditing. Featuring batch uploads and uncertainty flagging, Co-TA allows instructors to generate rubric variations and immediately simulate their impact against synthetic student personas. By visualizing resulting grade distributions prior to live grading, Co-TA serves as a critical fairness auditing mechanism, ensuring AI assessment remains transparent, verifiable, and firmly under human pedagogical control.

Cite this Paper


BibTeX
@InProceedings{pmlr-v339-jin26a, title = {Co-TA: Streamlining Automated Grading with Generative Rubrics and Distribution Simulation}, author = {Jin, Helen and Chen, Albert}, booktitle = {Proceedings of the Impactful and Responsible AI Systems for Education Workshop}, pages = {146--159}, year = {2026}, editor = {Basu Mallick, Debshila and Woodhead, Simon and Wang, Zichao and Ananda, Muktha and Burstein, Jill and Murphy, April}, volume = {339}, series = {Proceedings of Machine Learning Research}, month = {28 Jun}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v339/main/assets/jin26a/jin26a.pdf}, url = {https://proceedings.mlr.press/v339/jin26a.html}, abstract = {Large Language Models (LLMs) offer great potential to reduce Teaching Assistant workloads through automated grading. However, their rapid integration shifts the primary challenge from technical scalability to responsible governance, as LLMs frequently diverge from expert human graders on complex rubrics. Existing Human-in-the-Loop (HITL) systems focus on micro-level calibration for individual submissions, leaving instructors blind to how rubrics alter class-wide grade distributions. To address this gap, we introduce \textbf{Co-TA}, a deployable web application that streamlines the grading workflow while shifting the paradigm to rubric co-design and fairness auditing. Featuring batch uploads and uncertainty flagging, Co-TA allows instructors to generate rubric variations and immediately simulate their impact against synthetic student personas. By visualizing resulting grade distributions prior to live grading, Co-TA serves as a critical fairness auditing mechanism, ensuring AI assessment remains transparent, verifiable, and firmly under human pedagogical control.} }
Endnote
%0 Conference Paper %T Co-TA: Streamlining Automated Grading with Generative Rubrics and Distribution Simulation %A Helen Jin %A Albert Chen %B Proceedings of the Impactful and Responsible AI Systems for Education Workshop %C Proceedings of Machine Learning Research %D 2026 %E Debshila Basu Mallick %E Simon Woodhead %E Zichao Wang %E Muktha Ananda %E Jill Burstein %E April Murphy %F pmlr-v339-jin26a %I PMLR %P 146--159 %U https://proceedings.mlr.press/v339/jin26a.html %V 339 %X Large Language Models (LLMs) offer great potential to reduce Teaching Assistant workloads through automated grading. However, their rapid integration shifts the primary challenge from technical scalability to responsible governance, as LLMs frequently diverge from expert human graders on complex rubrics. Existing Human-in-the-Loop (HITL) systems focus on micro-level calibration for individual submissions, leaving instructors blind to how rubrics alter class-wide grade distributions. To address this gap, we introduce \textbf{Co-TA}, a deployable web application that streamlines the grading workflow while shifting the paradigm to rubric co-design and fairness auditing. Featuring batch uploads and uncertainty flagging, Co-TA allows instructors to generate rubric variations and immediately simulate their impact against synthetic student personas. By visualizing resulting grade distributions prior to live grading, Co-TA serves as a critical fairness auditing mechanism, ensuring AI assessment remains transparent, verifiable, and firmly under human pedagogical control.
APA
Jin, H. & Chen, A.. (2026). Co-TA: Streamlining Automated Grading with Generative Rubrics and Distribution Simulation. Proceedings of the Impactful and Responsible AI Systems for Education Workshop, in Proceedings of Machine Learning Research 339:146-159 Available from https://proceedings.mlr.press/v339/jin26a.html.

Related Material