2024/07/19 by Peng Cui, Vilém Zouhar, Cui, Peng +5
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Educational Methods and Media Use #FOS: Computer and information sciences #Reading and Literacy Development
paper · pdf · doi:10.48550/arxiv.2407.14309
openalex publication_date 2024/07/19 · openalex created_date 2024/09/26 · openalex updated_date 2026/07/28
Using questions in written text is an effective strategy to enhance readability. However, what makes an active reading question good, what the linguistic role of these questions is, and what is their impact on human reading remains understudied. We introduce GuidingQ, a dataset of 10K in-text questions from textbooks and scientific articles. By analyzing the dataset, we present a comprehensive understanding of the use, distribution, and linguistic characteristics of these questions. Then, we explore various approaches to generate such questions using language models. Our results highlight the importance of capturing inter-question relationships and the challenge of question position identification in generating these questions. Finally, we conduct a human study to understand the implication of such questions on reading comprehension. We find that the generated questions are of high quality and are almost as effective as human-written questions in terms of improving readers' memorization and comprehension.