2017/09/02 by Hassan Sajjad, Sajjad, Hassan, Fahim Dalvi +9
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1709.00616
openalex publication_date 2017/09/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Word segmentation plays a pivotal role in improving any Arabic NLP\napplication. Therefore, a lot of research has been spent in improving its\naccuracy. Off-the-shelf tools, however, are: i) complicated to use and ii)\ndomain/dialect dependent. We explore three language-independent alternatives to\nmorphological segmentation using: i) data-driven sub-word units, ii) characters\nas a unit of learning, and iii) word embeddings learned using a character CNN\n(Convolution Neural Network). On the tasks of Machine Translation and POS\ntagging, we found these methods to achieve close to, and occasionally surpass\nstate-of-the-art performance. In our analysis, we show that a neural machine\ntranslation system is sensitive to the ratio of source and target tokens, and a\nratio close to 1 or greater, gives optimal performance.\n