2024/08/21 by Tianyi Liu, Julia Chatain, Liu, Tianyi +9 · 6 citations
Computer Science · Engineering · Mathematics · #97-02 #Artificial intelligence #Computer science #Edcuational Technology Systems #Engineering #FOS: Mathematics #Grading (engineering) #History and Overview (math.HO) #Intelligent Tutoring Systems and Adaptive Learning #Mathematics #Mathematics education #Natural language processing #Online Learning and Analytics
paper · pdf · doi:10.48550/arxiv.2408.11728
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/08/21 · openalex created_date 2024/12/16 · openalex updated_date 2026/07/28
Effective and timely feedback in educational assessments is essential but labor-intensive, especially for complex tasks. Recent developments in automated feedback systems, ranging from deterministic response grading to the evaluation of semi-open and open-ended essays, have been facilitated by advances in machine learning. The emergence of pre-trained Large Language Models, such as GPT-4, offers promising new opportunities for efficiently processing diverse response types with minimal customization. This study evaluates the effectiveness of a pre-trained GPT-4 model in grading semi-open handwritten responses in a university-level mathematics exam. Our findings indicate that GPT-4 provides surprisingly reliable and cost-effective initial grading, subject to subsequent human verification. Future research should focus on refining grading rules and enhancing the extraction of handwritten responses to further leverage these technologies.