2018/09/11 by Yonatan Belinkov, Belinkov, Yonatan, Alexander Magidow +7 · 1 citation
Arts and Humanities · Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Language, Linguistics, Cultural Analysis #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1809.03891
openalex publication_date 2018/09/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Arabic is a widely-spoken language with a long and rich history, but existing\ncorpora and language technology focus mostly on modern Arabic and its\nvarieties. Therefore, studying the history of the language has so far been\nmostly limited to manual analyses on a small scale. In this work, we present a\nlarge-scale historical corpus of the written Arabic language, spanning 1400\nyears. We describe our efforts to clean and process this corpus using Arabic\nNLP tools, including the identification of reused text. We study the history of\nthe Arabic language using a novel automatic periodization algorithm, as well as\nother techniques. Our findings confirm the established division of written\nArabic into Modern Standard and Classical Arabic, and confirm other established\nperiodizations, while suggesting that written Arabic may be divisible into\nstill further periods of development.\n