vix.ing · top · new · best · stats · spec

Dotless Arabic Text for Natural Language Processing

2024/09/12 by Maged S. Al-Shaibani, Irfan Ahmad · 1 voice
Computer Science · #Advanced Computational Techniques and Applications #Handwritten Text Recognition Techniques #Natural Language Processing Techniques

paper · pdf · doi:10.1162/coli_a_00535

openalex publication_date 2024/09/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/03

Abstract

Abstract This article introduces a novel representation of Arabic text as an alternative approach for Arabic NLP, inspired by the dotless script of ancient Arabic. We explored this representation through extensive analysis on various text corpora, differing in size and domain, and tokenized using multiple tokenization techniques. Furthermore, we examined the information density of this representation and compared it with the standard dotted Arabic text using text entropy analysis. Utilizing parallel corpora, we also drew comparisons between Arabic and English text analysis to gain additional insights. Our investigation extended to various upstream and downstream NLP tasks, including language modeling, text classification, sequence labeling, and machine translation, examining the implications of both the representations. Specifically, we performed seven different downstream tasks using various tokenization schemes comparing the standard dotted text with dotless Arabic text representations. Performance using both the representations was comparable across different tokenizations. However, dotless representation achieves these results with significant reduction in vocabulary sizes, and in some scenarios showing reduction of up to 50%. Additionally, we present a system that restores dots to the dotless Arabic text. This system is useful for tasks that require Arabic texts as output.

Citations

Discussions

Related