2025/01/21 by Noah L. Schroeder, Chris Davis Jaldi, Schroeder, Noah L. +5 · 1 citation
Computer Science · #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Topic Modeling #cs.HC
paper · pdf · doi:10.48550/arxiv.2501.11840
openalex publication_date 2025/01/21 · openalex created_date 2025/10/10 · arxiv created 2026/07/29 · arxiv updated 2026/07/31 · openalex updated_date 2026/08/03
Systematic reviews are time-consuming endeavors that require knowledgeable human reviewers to screen studies for relevance and extract data following a specific coding scheme before any analysis or synthesis can occur. Large language models (LLMs) hold promise for substantially accelerating this process and reducing reviewer workload, yet their application within the context of systematic reviews in the field of education remains underexplored. We address this issue in two ways: through empirical studies and the iterative development of an open-source software tool. First, we conducted two empirical studies examining the efficacy of using LLMs for data extraction using data from a published review on pedagogical agents. We extracted a variety of data types from 112 studies and compared the results to data extracted by human coding. Results indicate that LLMs struggled with extracting data accurately and therefore are not ready to be used as primary data extraction tools without explicit human validation of the data extracted. These findings highlight the dire need for a human-in-the-loop (HIL) approach to AI-assisted data extraction. We then propose a HIL workflow and introduce and describe the development of a free, web-based, open-source tool designed to support user-friendly, human-validated data extraction with LLMs.