2018/06/10 by Antonis Anastasopoulos, Marika Lekakou, Anastasopoulos, Antonis +9 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1806.03757
openalex publication_date 2018/06/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Most work on part-of-speech (POS) tagging is focused on high resource\nlanguages, or examines low-resource and active learning settings through\nsimulated studies. We evaluate POS tagging techniques on an actual endangered\nlanguage, Griko. We present a resource that contains 114 narratives in Griko,\nalong with sentence-level translations in Italian, and provides gold\nannotations for the test set. Based on a previously collected small corpus, we\ninvestigate several traditional methods, as well as methods that take advantage\nof monolingual data or project cross-lingual POS tags. We show that the\ncombination of a semi-supervised method with cross-lingual transfer is more\nappropriate for this extremely challenging setting, with the best tagger\nachieving an accuracy of 72.9%. With an applied active learning scheme, which\nwe use to collect sentence-level annotations over the test set, we achieve\nimprovements of more than 21 percentage points.\n