2020/06/20 by Mahmoud Daif, Daif, Mahmoud, Shunsuke Kitada +3
Computer Science · #Text and Document Classification Technologies #Handwritten Text Recognition Techniques #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2006.11586
Classical and some deep learning techniques for Arabic text classification\noften depend on complex morphological analysis, word segmentation, and\nhand-crafted feature engineering. These could be eliminated by using\ncharacter-level features. We propose a novel end-to-end Arabic document\nclassification framework, Arabic document image-based classifier (AraDIC),\ninspired by the work on image-based character embeddings. AraDIC consists of an\nimage-based character encoder and a classifier. They are trained in an\nend-to-end fashion using the class balanced loss to deal with the long-tailed\ndata distribution problem. To evaluate the effectiveness of AraDIC, we created\nand published two datasets, the Arabic Wikipedia title (AWT) dataset and the\nArabic poetry (AraP) dataset. To the best of our knowledge, this is the first\nimage-based character embedding framework addressing the problem of Arabic text\nclassification. We also present the first deep learning-based text classifier\nwidely evaluated on modern standard Arabic, colloquial Arabic and classical\nArabic. AraDIC shows performance improvement over classical and deep learning\nbaselines by 12.29% and 23.05% for the micro and macro F-score, respectively.\n