vix.ing · top · new · best · stats · spec

Grammatical Error Correction for Low-Resource Languages: The Case of Zarma

2024/10/20 by Mamadou Keïta, Keita, Mamadou K., Adwoa Bremang +10 · 2 citations
Arts and Humanities · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Language, Linguistics, Cultural Analysis #Machine Learning (cs.LG)

paper · pdf · doi:10.48550/arxiv.2410.15539

openalex publication_date 2024/10/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Grammatical error correction (GEC) aims to improve text quality and readability. Previous work on the task focused primarily on high-resource languages, while low-resource languages lack robust tools. To address this shortcoming, we present a study on GEC for Zarma, a language spoken by over five million people in West Africa. We compare three approaches: rule-based methods, machine translation (MT) models, and large language models (LLMs). We evaluated GEC models using a dataset of more than 250,000 examples, including synthetic and human-annotated data. Our results showed that the MT-based approach using M2M100 outperforms others, with a detection rate of 95.82% and a suggestion accuracy of 78.90% in automatic evaluations (AE) and an average score of 3.0 out of 5.0 in manual evaluation (ME) from native speakers for grammar and logical corrections. The rule-based method was effective for spelling errors but failed on complex context-level errors. LLMs -- Gemma 2b and MT5-small -- showed moderate performance. Our work supports use of MT models to enhance GEC in low-resource settings, and we validated these results with Bambara, another West African language.

Cited by

Related