English

Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams

Computers and Society 2024-12-13 v2 Computation and Language Machine Learning

Abstract

Recent advances in generative artificial intelligence (AI) have shown promise in accurately grading open-ended student responses. However, few prior works have explored grading handwritten responses due to a lack of data and the challenge of combining visual and textual information. In this work, we leverage state-of-the-art multi-modal AI models, in particular GPT-4o, to automatically grade handwritten responses to college-level math exams. Using real student responses to questions in a probability theory exam, we evaluate GPT-4o's alignment with ground-truth scores from human graders using various prompting techniques. We find that while providing rubrics improves alignment, the model's overall accuracy is still too low for real-world settings, showing there is significant room for growth in this task.

Keywords

Cite

@article{arxiv.2411.05231,
  title  = {Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams},
  author = {Adriana Caraeni and Alexander Scarlatos and Andrew Lan},
  journal= {arXiv preprint arXiv:2411.05231},
  year   = {2024}
}

Comments

Published in LAK 2025: The 15th International Learning Analytics and Knowledge Conference

R2 v1 2026-06-28T19:52:28.261Z