MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
2024/10/24 by S Sakshi, Utkarsh Tyagi, Sakshi, S +15 · 93 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2410.19168
Abstract
The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges posed by MMAU. Notably, even the most advanced Gemini Pro v1.5 achieves only 52.97% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.
Cited by
- Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
- Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating
- Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- Protecting Bystander Privacy via Selective Hearing in Audio LLMs
- Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
- ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- Step-Audio-R1 Technical Report
- Multimodal Evaluation of Russian-language Architectures
- Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
- Spatial Blind Spot: Auditory Motion Perception Deficits in Audio LLMs
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- SAR-LM: Symbolic Audio Reasoning with Large Language Models
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
- Assessing Factual Music Comprehension in Large Audio Language Models
- LongCat-Flash-Omni Technical Report
- NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
- Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- Large language model-based task planning for service robots: A review
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- Evaluating Multimodal Large Language Models on Core Music Perception Tasks
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
- M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
- The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
- Can large audio language models understand child stuttering speech? speech summarization, and source separation
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
- UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
- Scaling Language-Centric Omnimodal Representation Learning
- Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
- MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval
- AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
- Robustness assessment of large audio language models in multiple-choice evaluation
- AURA Score: A Metric For Holistic Audio Question Answering Evaluation
- AudioToolAgent: An Agentic Framework for Audio-Language Models
- When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
- Hearing the Order: Investigating Selection Bias in Large Audio-Language Models
- When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models
- PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
- TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
- Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
- MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark
- Investigating Faithfulness in Large Audio Language Models
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models
- Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
- Investigating Modality Contribution in Audio LLMs for Music
- Benchmarking Gaslighting Attacks Against Speech Large Language Models
- Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
- Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
- AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
- Advancing Speech Understanding in Speech-Aware Language Models with GRPO
- AudioGenie-Reasoner: A Training-Free Multi-Agent Framework for Coarse-to-Fine Audio Deep Reasoning
- FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
- Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
- SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
- SAM: A Mamba-2 State-Space Audio-Language Model
- Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
- Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering
- Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
- AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- Advancing the Foundation Model for Music Understanding
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
- DIFFA: Large Language Diffusion Models Can Listen and Understand
- TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
- ERNIE 5.0 Technical Report
- Step-Audio 2 Technical Report
- AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
- MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
Related