2025/03/31 by Ngoc Dung Huynh, Mohamed Reda Bouadjenek, Huynh, Ngoc Dung +7
Computer Science · #FOS: Computer and information sciences #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Speech and dialogue systems #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2503.24164
openalex publication_date 2025/03/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Large vision and language models show strong performance in tasks like image captioning, visual question answering, and retrieval. However, challenges remain in integrating speech, text, and vision into a unified model, especially for spoken tasks. Speech generation methods vary (some produce speech directly), others through text (but their impact on quality is unclear). Evaluation often relies on automatic speech recognition, which may introduce bias. We propose SVLA, a unified speech vision language model based on a transformer architecture that handles multimodal inputs and outputs. We train it on 38.2 million speech text image examples, including 64.1 hours of synthetic speech. We also introduce Speech VQA Accuracy, a new metric for evaluating spoken responses. SVLA improves multimodal understanding and generation by better combining speech, vision, and language.