vix.ing · top · new · best · stats · spec

Towards Multi-Modal Mastery: A 4.5B Parameter Truly Multi-Modal Small Language Model

2024/11/08 by Ben Koska, Koska, Ben, Mojmír Horváth +1
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2411.05903

openalex publication_date 2024/11/08 · openalex created_date 2024/11/15 · openalex updated_date 2026/07/28

Abstract

We present a novel 4.5B parameter small language model that can handle multiple input and output modalities, including text, images, videos, and audio. Despite its small size, the model achieves near state-of-the-art performance on a variety of tasks, demonstrating the potential of multi-modal models to tackle complex real-world problems. Our approach leverages recent advancements in language modeling and multi-task learning to create a versatile and high-performing model that can even be deployed for edge inference. Experimental results show the model's strong performance across multiple benchmarks, paving the way for further progress in multi-modal artificial intelligence.

Related