2025/08/20 by Mukhammadsaid Mamasaidov, Azizullah Aral, Mamasaidov, Mukhammadsaid +5
Economics, Econometrics and Finance · Social Sciences · #Central Asia Education and Culture #Computation and Language (cs.CL) #Economic and Industrial Development #Education, Innovation and Language Studies #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2508.14586
openalex publication_date 2025/08/20 · openalex created_date 2025/10/16 · openalex updated_date 2026/07/28
Southern Uzbek (uzs) is a Turkic language variety spoken by around 5 million people in Afghanistan and differs significantly from Northern Uzbek (uzn) in phonology, lexicon, and orthography. Despite the large number of speakers, Southern Uzbek is underrepresented in natural language processing. We present new resources for Southern Uzbek machine translation, including a 997-sentence FLORES+ dev set, 39,994 parallel sentences from dictionary, literary, and web sources, and a fine-tuned NLLB-200 model (lutfiy). We also propose a post-processing method for restoring Arabic-script half-space characters, which improves handling of morphological boundaries. All datasets, models, and tools are released publicly to support future work on Southern Uzbek and other low-resource languages.