vix.ing · top · new · best · stats · spec

An Analysis of Language Frequency and Error Correction for Esperanto

2024/02/15 by Junhong Liang, Liang, Junhong
Arts and Humanities · Psychology · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Linguistics and Cultural Studies #Linguistics, Language Diversity, and Identity #Phonetics and Phonology Research

paper · pdf · doi:10.48550/arxiv.2402.09696

openalex publication_date 2024/02/15 · openalex created_date 2024/02/18 · openalex updated_date 2026/07/28

Abstract

Current Grammar Error Correction (GEC) initiatives tend to focus on major languages, with less attention given to low-resource languages like Esperanto. In this article, we begin to bridge this gap by first conducting a comprehensive frequency analysis using the Eo-GP dataset, created explicitly for this purpose. We then introduce the Eo-GEC dataset, derived from authentic user cases and annotated with fine-grained linguistic details for error identification. Leveraging GPT-3.5 and GPT-4, our experiments show that GPT-4 outperforms GPT-3.5 in both automated and human evaluations, highlighting its efficacy in addressing Esperanto's grammatical peculiarities and illustrating the potential of advanced language models to enhance GEC strategies for less commonly studied languages.

Related