2025/05/13 by N. M. Alam, Alam, Nahid, Karthik Reddy Kanjula +34 · 3 citations
Arts and Humanities · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Media, Religion, Digital Communication
paper · pdf · doi:10.48550/arxiv.2505.08910
openalex publication_date 2025/05/13 · openalex created_date 2025/10/15 · openalex updated_date 2026/07/28
In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily in widely spoken languages but lack performance on low-resource languages and varied cultural contexts. To address these limitations, we introduce Maya, an open-source Multilingual VLM. Our contributions are: 1) a multilingual image-text pretraining dataset in eight languages, based on the LLaVA pretraining dataset; and 2) a multilingual image-text model supporting these languages, enhancing cultural and linguistic comprehension in vision-language tasks. Code available at https://github.com/nahidalam/maya.