2016/03/06 by Giuliano Lancioni, Lancioni, Giuliano, Valeria Pettinari +11
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Natural Language Processing Techniques #Text and Document Classification Technologies #Topic Modeling #cs.CL #cs.IR
paper · pdf · doi:10.48550/arxiv.1603.01833
arxiv created 2016/03/06 · openalex publication_date 2016/03/06 · arxiv updated 2016/03/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
An extended, revised form of Tim Buckwalter's Arabic lexical and morphological resource AraMorph, eXtended Revised AraMorph (henceforth XRAM), is presented which addresses a number of weaknesses and inconsistencies of the original model by allowing a wider coverage of real-world Classical and contemporary (both formal and informal) Arabic texts. Building upon previous research, XRAM enhancements include (i) flag-selectable usage markers, (ii) probabilistic mildly context-sensitive POS tagging, filtering, disambiguation and ranking of alternative morphological analyses, (iii) semi-automatic increment of lexical coverage through extraction of lexical and morphological information from existing lexical resources. Testing of XRAM through a front-end Python module showed a remarkable success level.