2025/04/07 by Toqeer Ehsan, Ehsan, Toqeer, Thamar Solorio +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Network Packet Processing and Optimization #Neural Networks and Applications #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2504.08792
openalex publication_date 2025/04/07 · openalex created_date 2025/10/14 · openalex updated_date 2026/07/28
Named Entity Recognition (NER), a fundamental task in Natural Language Processing (NLP), has shown significant advancements for high-resource languages. However, due to a lack of annotated datasets and limited representation in Pre-trained Language Models (PLMs), it remains understudied and challenging for low-resource languages. To address these challenges, we propose a data augmentation technique that generates culturally plausible sentences and experiments on four low-resource Pakistani languages; Urdu, Shahmukhi, Sindhi, and Pashto. By fine-tuning multilingual masked Large Language Models (LLMs), our approach demonstrates significant improvements in NER performance for Shahmukhi and Pashto. We further explore the capability of generative LLMs for NER and data augmentation using few-shot learning.