2023/09/27 by Seungwhan Moon, Andrea Madotto, Moon, Seungwhan +24 · 1 voice · 11 citations
Computer Science · #Artificial intelligence #Computer science #Cover (algebra) #Human–computer interaction #Modality (human–computer interaction) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Programming language #Scalability #Set (abstract data type) #Space (punctuation) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2309.16058
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/09/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present Any-Modality Augmented Language Model (AnyMAL), a unified model that reasons over diverse input modality signals (i.e. text, image, video, audio, IMU motion sensor), and generates textual responses. AnyMAL inherits the powerful text-based reasoning abilities of the state-of-the-art LLMs including LLaMA-2 (70B), and converts modality-specific signals to the joint textual space through a pre-trained aligner module. To further strengthen the multimodal LLM's capabilities, we fine-tune the model with a multimodal instruction set manually collected to cover diverse topics and tasks beyond simple QAs. We conduct comprehensive empirical analysis comprising both human and automatic evaluations, and demonstrate state-of-the-art performance on various multimodal tasks.