2025/12/05 by Iris Schepers, L.M. Bruijn, Martijn Wieling +1 · 1 voice
Social Sciences · Medicine · Computer Science · #Artificial Intelligence in Law #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.1007/s10506-025-09495-1
openalex publication_date 2025/12/05 · openalex created_date 2025/12/05 · openalex updated_date 2026/07/23
Abstract This paper evaluates the performance of GPT-4o in annotating decisions of the United Nations Committee on Economic, Social, and Cultural Rights and compares these results with manual annotations by trained law students and senior (legal) scholars. GPT-4o achieves human-level accuracy in basic annotations, but struggles with recall in citation extraction, particularly for complex legal references. Human annotators, while more reliable in citation extraction, introduce formatting inconsistencies and occasional errors due to sloppiness. In contrast, GPT-4o maintains high precision, but suffers from variability across repeated prompts, raising concerns about reproducibility. Beyond accuracy, this study highlights cost-effectiveness as a key advantage of GPT-4o. The model significantly reduces annotation time and expenses compared to human annotators, who require post-processing and expert supervision. While GPT-4o produces structured output with fewer formatting inconsistencies, its omissions and inconsistencies require human oversight. These findings highlight trade-offs in expertise, cost, and reliability between human and AI-driven annotation. Although GPT-4o is a viable tool for basic legal annotations, improvements in recall and consistency are needed for more complex tasks.