Practical and ethical challenges of large language models in education: A systematic scoping review
2023/08/06 by Lixiang Yan, Lele Sha, Linxuan Zhao +7 · 57 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Online Learning and Analytics #Topic Modeling
paper · pdf · doi:10.1111/bjet.13370
openalex publication_date 2023/08/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Abstract Educational technology innovations leveraging large language models (LLMs) have shown the potential to automate the laborious process of generating and analysing textual content. While various innovations have been developed to automate a range of educational tasks (eg, question generation, feedback provision, and essay grading), there are concerns regarding the practicality and ethicality of these innovations. Such concerns may hinder future research and the adoption of LLMs‐based innovations in authentic educational contexts. To address this, we conducted a systematic scoping review of 118 peer‐reviewed papers published since 2017 to pinpoint the current state of research on using LLMs to automate and support educational tasks. The findings revealed 53 use cases for LLMs in automating education tasks, categorised into nine main categories: profiling/labelling, detection, grading, teaching support, prediction, knowledge representation, feedback, content generation, and recommendation. Additionally, we also identified several practical and ethical challenges, including low technological readiness, lack of replicability and transparency and insufficient privacy and beneficence considerations. The findings were summarised into three recommendations for future studies, including updating existing innovations with state‐of‐the‐art models (eg, GPT‐3/4), embracing the initiative of open‐sourcing models/systems, and adopting a human‐centred approach throughout the developmental process. As the intersection of AI and education is continuously evolving, the findings of this study can serve as an essential reference point for researchers, allowing them to leverage the strengths, learn from the limitations, and uncover potential research opportunities enabled by ChatGPT and other generative AI models. Practitioner notes What is currently known about this topic Generating and analysing text‐based content are time‐consuming and laborious tasks. Large language models are capable of efficiently analysing an unprecedented amount of textual content and completing complex natural language processing and generation tasks. Large language models have been increasingly used to develop educational technologies that aim to automate the generation and analysis of textual content, such as automated question generation and essay scoring. What this paper adds A comprehensive list of different educational tasks that could potentially benefit from LLMs‐based innovations through automation. A structured assessment of the practicality and ethicality of existing LLMs‐based innovations from seven important aspects using established frameworks. Three recommendations that could potentially support future studies to develop LLMs‐based innovations that are practical and ethical to implement in authentic educational contexts. Implications for practice and/or policy Updating existing innovations with state‐of‐the‐art models may further reduce the amount of manual effort required for adapting existing models to different educational tasks. The reporting standards of empirical research that aims to develop educational technologies using large language models need to be improved. Adopting a human‐centred approach throughout the developmental process could contribute to resolving the practical and ethical challenges of large language models in education.
Citations
Cited by
- Problems With Large Language Models for Learner Modelling: Why LLMs Alone Fall Short for Responsible Tutoring in K--12 Education
- A systematic review of generative AI in education: Empirical insights from a human– AI interaction perspective
- Trusting AI to detect AI? A systematic evaluation of the reliability and robustness of current AIGC detection tools for student academic work
- An Information-Theoretic Framework for Robust Large Language Model Editing
- Modeling Collaborative Problem Solving Dynamics from Group Discourse: A Text-Mining Approach with Synergy Degree Model
- Strategic Innovation Management in the Age of Large Language Models Market Intelligence, Adaptive R&D, and Ethical Governance
- Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
- Amplifier or substitute? A systematic review of generative AI’s impact on higher-order cognitive skills among university students
- The cognitive biases that may exacerbate inflationary and deflationary positions about large language models
- Transformación digital en la educación: percepciones estudiantiles sobre la incorporación de la Inteligencia Artificial
- Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
- AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence
- Ethical AI prompt recommendations in large language models using collaborative filtering
- Open-Source Large Language Models in Education: A Narrative Review of Evidence, Pedagogical Roles, and Learning Outcomes
- Conceptualizing How to Design for AI Literacy through Game Artifacts
- Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
- Mind the Ethics! The Overlooked Ethical Dimensions of GenAI in Software Modeling Education
- SCENIC: A Location-based System to Foster Cognitive Development in Children During Car Rides
- LearnLens: An AI-Enhanced Dashboard to Support Teachers in Open-Ended Classrooms
- LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
- RETRACTED ARTICLE: The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: insights from a meta-analysis
- Navigating the AI era: university communication strategies and perspectives on generative AI tools
- LLM Chatbot-Creation Approaches
- Impact of AI assistance on student agency
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- The potential and limitations of large language models for automatic classification of teachers' motivational messages in educational research
- A ChatGPT-based approach for questions generation in higher education
- Beyond the surface: stylometric analysis of GPT-4o’s capacity for literary style imitation
- Examining the consistency of instructor versus large language model ratings on summary content: Toward checklist-based feedback provision with second language writers
- Bridging MOOCs, Smart Teaching, and AI: A Decade of Evolution Toward a Unified Pedagogy
- AI-based research mentors: Plausible scenarios and ethical issues
- ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning
- A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots
- Can GPT-4o Evaluate Usability Like Human Experts? A Comparative Study on Issue Identification in Heuristic Evaluation
- Beyond the “Wow” Factor: Using Generative AI for Increasing Generative Sense-Making
- Position: EU AI Act's Research Exemptions Can Break the Publication Norms of Major AI Conferences
- Generative AI and English language teaching: A global Englishes perspective
- Supporting teachers' value‐sensitive reflections on the cost–benefit dynamics of technology in educational practices
- Evaluation of LLMs for mathematical problem solving
- Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data
- The promise and challenges of generative AI in education
- Uses of artificial intelligence and machine learning in systematic reviews of education research
- Examining the Role of LLM-Driven Interactions on Attention and Cognitive Engagement in Virtual Classrooms
- Highlight: MEGA into the New Generation of Computational Genetics
- Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
- Attack and defense techniques in large language models: A survey and new perspectives
- Generative AI in Higher Education Laboratory Learning: A Qualitative Case Study of Epistemic Scaffolding and Assessment Boundaries
- Descriptive writing using generative AI as a cognitive scaffold in the metaverse environment: university students’ perceptions, learning engagement, and performance
- Automated Grading of Students’ Short Answers Using Language Models
- Stan: An LLM-based thermodynamics course assistant
- The effects of generative AI agents and scaffolding on enhancing students’ comprehension of visual learning analytics
- Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning
- Assessing novice programmers' perception of ChatGPT:performance, risk, decision-making, and intentions
- Towards responsible AI for education: Hybrid human-AI to confront the Elephant in the room
- Integrating LLMs for Grading and Appeal Resolution in Computer Science Education
- A Design-Based Research Approach to What Distance Learners Expect and Value from an Institutional AI Digital Assistant
- Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index
Related