2026/03/31 by Marcos Abreu, Alvaro Suárez, Álvaro Suárez +3
Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Intelligent Tutoring Systems and Adaptive Learning
paper · doi:10.1088/1361-6552/ae8110
Abstract Recent advances in large language models have opened new possibilities for Artificial Intelligence (AI)-assisted grading in higher education. This study explores the use of a GPT-5.4–based system to support the evaluation of introductory physics laboratory reports using a rubric-driven automated workflow implemented in batch mode via API. AI-generated scores were compared with instructor grading, and the feedback was analysed qualitatively across rubric criteria. The results show weak agreement in the ranking of reports and noticeable differences at the level of individual scores. Item-level analysis indicates that the model frequently produced feedback classified as correct application but also generated a non-negligible proportion of reasonable but superficial responses and invalid evaluations. These limitations were mainly associated with restricted access to evidence during text extraction and OCR, particularly for equations, figures, and graphical representations. The findings suggest that batch-based AI grading can provide systematic feedback and help identify recurring patterns in student work, supporting instructors in large-scale courses. However, under the conditions examined, AI-generated scores were not interchangeable with instructor grading and should be understood as a tool to assist, rather than replace, the evaluation process.