From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data
2025/05/26 by Kuan, Chun-Yi, Lee, Hung-yi · 1 citation
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2505.20166
Abstract
Audio-aware large language models (ALLMs) have recently made great strides in understanding and processing audio inputs. These models are typically adapted from text-based large language models (LLMs) through additional training on audio-related tasks. However, this adaptation process presents two major limitations. First, ALLMs often suffer from catastrophic forgetting, where crucial textual capabilities like instruction-following are lost after training on audio data. In some cases, models may even hallucinate sounds that are not present in the input audio, raising concerns about reliability. Second, achieving cross-modal alignment between audio and language typically relies on large collections of task-specific question-answer pairs for instruction tuning, making it resource-intensive. To address these issues, previous works have leveraged the backbone LLMs to synthesize general-purpose, caption-style alignment data. In this paper, we propose a data generation framework that produces contrastive-like training data, designed to enhance ALLMs' ability to differentiate between present and absent sounds. We further extend our approach to multi-audio scenarios, enabling the model to either explain differences between audio inputs or produce unified captions that describe all inputs, thereby enhancing audio-language alignment. We refer to the entire ALLM training framework as bootstrapping audio-language alignment via synthetic data generation from backbone LLMs (BALSa). Experimental results indicate that our method effectively mitigates audio hallucinations while reliably maintaining strong performance on audio understanding and reasoning benchmarks, as well as instruction-following skills. Moreover, incorporating multi-audio training further enhances the model's comprehension and reasoning capabilities. Overall, BALSa offers an efficient and scalable approach to developing ALLMs.
Citations
- Step-Audio 2 Technical Report
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Discrete Audio Tokens: More Than a Survey!
- Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
- Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
- SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
- Qwen3 Technical Report
- Kimi-Audio Technical Report
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- Qwen2.5-Omni Technical Report
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
- Recent Advances in Discrete Speech Tokens: A Review
- ADIFF: Explaining audio difference using natural language
- Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
- Qwen2.5 Technical Report
- WavChat: A Survey of Spoken Dialogue Models
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
- A Survey on Speech Large Language Models for Understanding
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- Distilling an End-to-End Voice Assistant Without Instruction Training Data
- Recent Advances in Speech Language Models: A Survey
- DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
- The Llama 3 Herd of Models
- Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
- Qwen2-Audio Technical Report
- Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
- DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
- BLSP-Emo: Towards Empathetic Large Speech-Language Models
- Audio Dialogues: Dialogues dataset for audio and music understanding
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Towards audio language modeling -- an overview
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
- Joint Audio and Speech Understanding
- Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
- BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Pengi: An Audio Language Model for Audio Tasks
- Listen, Think, and Understand
- Evaluating Object Hallucination in Large Vision-Language Models
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Instruction Tuning with GPT-4
- A Survey of Large Language Models
- GPT-4 Technical Report
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Robust Speech Recognition via Large-Scale Weak Supervision
- CochlScene: Acquisition of acoustic scene data using crowdsourcing
- Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
- FSD50K: An Open Dataset of Human-Labeled Sound Events
- Language Models are Few-Shot Learners
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
Cited by
Related