2014/10/17 by Yannis Haralambous, Haralambous, Yannis, Yassir Elidrissi +3
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
paper · pdf · doi:10.48550/arxiv.1410.4863
10 pages, 4 figure, accepted at CITALA 2014 (http://www.citala.org/)
arxiv created 2014/10/17 · arxiv updated 2014/10/21
We study the performance of Arabic text classification combining various techniques: (a) tfidf vs. dependency syntax, for feature selection and weighting; (b) class association rules vs. support vector machines, for classification. The Arabic text is used in two forms: rootified and lightly stemmed. The results we obtain show that lightly stemmed text leads to better performance than rootified text; that class association rules are better suited for small feature sets obtained by dependency syntax constraints; and, finally, that support vector machines are better suited for large feature sets based on morphological feature selection criteria.