A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
2026/05/18 by Kaiwen Luo, Zhenhong Zhou, Leo Wang +34
Computer Science · #Adversarial Robustness in Machine Learning #Music and Audio Processing #Speech Recognition and Synthesis #cs.SD
paper · pdf · doi:10.48550/arxiv.2605.20266
openalex publication_date 2026/05/18 · openalex created_date 2026/05/22 · openalex updated_date 2026/07/28 · arxiv created 2026/08/03 · arxiv updated 2026/08/04
Abstract
The foundational capabilities established by Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs), within which Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite their remarkable performance, the escalation of LALMs' capabilities has significantly outpaced the development of systemic frameworks to ensure their trustworthiness. This survey provides a comprehensive investigation into the endogenous mechanisms of LALMs, detailing the architectural innovations and alignment algorithms that facilitate emergent reasoning. Specifically, we analyze how the transition to unified end-to-end frameworks and the integration of continuous acoustic signals inherently expand the attack surface. To rigorously evaluate the risks within these paradigms, we establish a comprehensive taxonomy of trustworthiness, categorizing critical vulnerabilities such as cross-modal jailbreaking, latent acoustic backdoors, and biometric privacy leakage. We review the state-of-the-art through six analytical pillars: hallucination, robustness, safety, privacy, fairness, and authentication. The profound imbalance between a mature offensive landscape and underdeveloped defenses further validates the critical trustworthiness gaps and multidimensional risks facing audio-centric intelligence. Finally, we propose a strategic roadmap advocating for "Defense-in-Depth" architectures, causal auditory world modeling, and intrinsic representation engineering to bridge the gap between empirical performance and intrinsically trustworthy audio intelligence. Our project has been uploaded to GitHub https://github.com/Kwwwww74/Awesome-Trustworthy-AudioLLMs.
Citations
- AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
- Qwen3.5-Omni Technical Report
- Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
- Qwen3-ASR Technical Report
- X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
- Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
- BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
- DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
- Towards Audio Token Compression in Large Audio Language Models
- It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
- Step-Audio-R1 Technical Report
- Segmentwise Pruning in Audio-Language Models
- Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models
- Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful Audio
- End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
- StyleBreak: Revealing Alignment Vulnerabilities in Large Audio-Language Models via Style-Aware Audio Jailbreak
- SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
- MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
- Step-Audio-EditX Technical Report
- SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
- Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
- The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- Can Speech LLMs Think while Listening?
- Sci-Phi: A Large Language Model Spatial Audio Descriptor
- Latent Speech-Text Transformer
- Robustness assessment of large audio language models in multiple-choice evaluation
- AudioToolAgent: An Agentic Framework for Audio-Language Models
- When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models
- Hearing the Order: Investigating Selection Bias in Large Audio-Language Models
- EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
- Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models
- VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- Investigating Faithfulness in Large Audio Language Models
- Investigating Modality Contribution in Audio LLMs for Music
- Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
- Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
- Benchmarking Gaslighting Attacks Against Speech Large Language Models
- Can Audio Large Language Models Verify Speaker Identity?
- Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
- Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
- SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
- EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
- From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
- Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
- FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
- When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
- Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
- When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
- Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers
- SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
- Step-Audio 2 Technical Report
- SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
- Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World
- DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
- WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
- SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models
- PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
- Fine-Tuning Large Audio-Language Models with LoRA for Precise Temporal Localization of Prolonged Exposure Therapy Elements
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
- SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
- Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Study
- Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
- Towards Reliable Large Audio Language Model
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
- VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
- SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
- AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
- SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
- VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
- Kimi-Audio Technical Report
- A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
- Multilingual and Multi-Accent Jailbreaking of Audio LLMs
- A Survey on Trustworthy LLM Agents: Threats and Countermeasures
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
- Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
- URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
- Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
- Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
- Large Language Model Safety: A Holistic Survey
- A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
- WavChat: A Survey of Spoken Dialogue Models
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- A Survey on Speech Large Language Models for Understanding
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
- Distilling an End-to-End Voice Assistant Without Instruction Training Data
- Recent Advances in Speech Language Models: A Survey
- LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
- A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
- Moshi: a speech-text foundation model for real-time dialogue
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- Qwen2-Audio Technical Report
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
- AudioBench: A Universal Benchmark for Audio Large Language Models
- Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- SpeechVerse: A Large-scale Generalizable Audio Language Model
- A Survey on Speech Deepfake Detection
- WavLLM: Towards Robust and Adaptive Speech Large Language Model
- Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations
- Spirit LM: Interleaved Spoken and Written Language Model
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
- SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
- E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
- Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
- SLM: Bridge the thin gap between speech and text foundation models
- Qwen Technical Report
- Joint Audio and Speech Understanding
- Audio Deepfake Detection: A Survey
- Sparks of Large Audio Models: A Survey and Outlook
- Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning
- Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
- Pengi: An Audio Language Model for Audio Tasks
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Generative Spoken Dialogue Language Modeling
- Training language models to follow instructions with human feedback
- PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
- Audio-Language Models for Audio-Centric Tasks: A Systematic Survey