vix.ing · top · new · best · stats

Benchmarking Large Language Models for Personalized Guidance in AI-Enhanced Learning

2025/09/02 by Bo Yuan, Jun Hu, Yuan, Bo +2 · 1 voice
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning #Topic Modeling #cs.AI

paper · pdf · doi:10.48550/arxiv.2509.05346

openalex publication_date 2025/09/02 · arxiv published 2025/09/02 · openalex created_date 2025/10/11 · arxiv updated 2025/10/22 · openalex updated_date 2026/07/28

Abstract

While Large Language Models (LLMs) are increasingly envisioned as intelligent assistants for personalized learning, systematic head-to-head evaluations in authentic learning scenarios remain scarce. This study presents an empirical comparison of three state-of-the-art LLMs on a tutoring task simulating a realistic learning setting. Using a dataset containing a student's responses to ten mixed-format questions with correctness labels, each model was asked to (i) analyze the quiz to identify underlying knowledge components, (ii) infer the student's mastery profile, and (iii) generate targeted guidance for improvement. To mitigate subjectivity and evaluator bias, Gemini was employed as a virtual judge to perform pairwise comparisons across multiple dimensions: accuracy, clarity, actionability, and appropriateness. Results analyzed via the Bradley-Terry model reveal that GPT-4o is generally preferred, producing feedback that is more informative and better structured than its counterparts, whereas DeepSeek-V3 and GLM-4.5 demonstrate intermittent strengths but lower consistency. These findings highlight the feasibility of deploying LLMs as advanced teaching assistants for individualized support and provide methodological insights for subsequent empirical research on LLM-driven personalized learning.

Citations

Discussions

Related