2024/06/21 by Subhankar Maity, Maity, Subhankar, Aniket Deroy +3 · 2 citations
Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Educational Assessment and Pedagogy #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2406.15211
openalex publication_date 2024/06/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We evaluate the effectiveness of GPT-4 Turbo in generating educational questions from NCERT textbooks in zero-shot mode. Our study highlights GPT-4 Turbo's ability to generate questions that require higher-order thinking skills, especially at the "understanding" level according to Bloom's Revised Taxonomy. While we find a notable consistency between questions generated by GPT-4 Turbo and those assessed by humans in terms of complexity, there are occasional differences. Our evaluation also uncovers variations in how humans and machines evaluate question quality, with a trend inversely related to Bloom's Revised Taxonomy levels. These findings suggest that while GPT-4 Turbo is a promising tool for educational question generation, its efficacy varies across different cognitive levels, indicating a need for further refinement to fully meet educational standards.