2026/02/13 by Adib Sakhawat, Shamim Ara Parveen, Md Ruhul Amin +4
Arts and Humanities · Psychology · #Bengali #Categorization, perception, and language #Language, Metaphor, and Cognition #Literal and figurative language #Metaphor #Paraphrase #Syntax, Semantics, Linguistic Variation
paper · pdf · doi:10.63317/546w2cys6m6t
openalex publication_date 2026/04/30 · openalex created_date 2026/05/05 · openalex updated_date 2026/08/05
Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grounded corpus of 10,361 Bengali idioms. Each idiom is annotated under a comprehensive 19-field schema, established and refined through a deliberative expert consensus process, that captures its semantic, syntactic, cultural, and religious dimensions, providing a rich, structured resource for computational linguistics. To establish a robust benchmark for Bangla figurative language understanding, we evaluate 30 state-of-the-art multilingual and instruction-tuned LLMs on the task of inferring figurative meaning. Our results reveal a critical performance gap, with no model surpassing 50% accuracy, a stark contrast to significantly higher human performance (83.4%). This underscores the limitations of existing models in cross-linguistic and cultural reasoning. By releasing the new idiom dataset and benchmark, we provide foundational infrastructure for advancing figurative language understanding and cultural grounding in LLMs for Bengali and other low-resource languages.