vix.ing · top · new · best · stats · spec

Shamela: A Large-Scale Historical Arabic Corpus

2016/12/28 by Yonatan Belinkov, Alexander Magidow, Belinkov, Yonatan +7 · 4 citations
Arts and Humanities · Business, Management and Accounting · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Islamic Finance and Banking Studies #Islamic Finance and Communication #Language, Linguistics, Cultural Analysis

paper · pdf · doi:10.48550/arxiv.1612.08989

openalex publication_date 2016/12/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Arabic is a widely-spoken language with a rich and long history spanning more than fourteen centuries. Yet existing Arabic corpora largely focus on the modern period or lack sufficient diachronic information. We develop a large-scale, historical corpus of Arabic of about 1 billion words from diverse periods of time. We clean this corpus, process it with a morphological analyzer, and enhance it by detecting parallel passages and automatically dating undated texts. We demonstrate its utility with selected case-studies in which we show its application to the digital humanities.

Cited by

Related