2018/08/23 by Wajdi Zaghouani, Zaghouani, Wajdi, Anis Charfi +1 · 3 citations
Computer Science · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #Digital Communication and Language #FOS: Computer and information sciences #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.1808.07674
openalex publication_date 2018/08/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we present Arap-Tweet, which is a large-scale and\nmulti-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab\nworld representing the major Arabic dialectal varieties. To build this corpus,\nwe collected data from Twitter and we provided a team of experienced annotators\nwith annotation guidelines that they used to annotate the corpus for age\ncategories, gender, and dialectal variety. During the data collection effort,\nwe based our search on distinctive keywords that are specific to the different\nArabic dialects and we also validated the location using Twitter API. In this\npaper, we report on the corpus data collection and annotation efforts. We also\npresent some issues that we encountered during these phases. Then, we present\nthe results of the evaluation performed to ensure the consistency of the\nannotation. The provided corpus will enrich the limited set of available\nlanguage resources for Arabic and will be an invaluable enabler for developing\nauthor profiling tools and NLP tools for Arabic.\n