2002/05/21 by Mathias Creutz, Creutz, Mathias, Krista Lagus +1 · 3 citations
Computer Science · #Algorithms and Data Compression #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Natural Language Processing Techniques #Web Data Mining and Analysis #cs.CL
paper · pdf · doi:10.48550/arxiv.cs/0205057
10 pages, to appear in Proceedings of Morphological and Phonological Learning Workshop of ACL'02
arxiv created 2002/05/21 · openalex publication_date 2002/05/21 · arxiv updated 2009/11/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present two methods for unsupervised segmentation of words into morpheme-like units. The model utilized is especially suited for languages with a rich morphology, such as Finnish. The first method is based on the Minimum Description Length (MDL) principle and works online. In the second method, Maximum Likelihood (ML) optimization is used. The quality of the segmentations is measured using an evaluation method that compares the segmentations produced to an existing morphological analysis. Experiments on both Finnish and English corpora show that the presented methods perform well compared to a current state-of-the-art system.