Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
2025/11/20 by Tseng, Wei-Cheng, Zhou, Xuanru, Huo, Mingyue +3
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2511.16757
Abstract
Audio-language pretraining holds promise for general-purpose audio understanding, yet remains underexplored compared to its vision counterpart. While vision-language models like CLIP serve as widely adopted foundations, existing audio-language models primarily excel at retrieval tasks with limited adoption as general-purpose encoders. We identify three key barriers: limited large-scale audio-text corpora, insufficient caption diversity, and lack of systematic exploration and evaluation. To this end, we introduce CaptionStew, a 10.7M caption dataset aggregating diverse open-source audio-text corpora across multiple domains and captioning styles. Using this resource, we conduct the first comprehensive evaluation comparing contrastive and captioning objectives for audio representation learning across speech, music, and environmental sound tasks. Our results demonstrate that audio-language pretraining yields competitive, transferable representations. Through systematic data-scaling experiments, we reveal complementary objective strengths: contrastive learning achieves superior data efficiency at smaller scales, while captioning demonstrates better scalability on language-involved audio understanding tasks. We also find that common supervised initialization practices provide diminishing returns at scale, challenging current approaches. These findings establish audio-language pretraining as a viable pathway toward general-purpose audio representations, guiding future research. To accelerate progress, we release data preparation recipes, training protocols, and pretrained models, paving the way toward universal audio understanding.
Citations
- Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
- Scaling Rich Style-Prompted Text-to-Speech Datasets
- JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
- MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
- Qwen2.5 Technical Report
- AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
- Multimodal Autoregressive Pre-training of Large Vision Encoders
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- Qwen2-Audio Technical Report
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- Scaling up masked audio encoder learning for general audio classification
- EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
- M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
- Cacophony: An Improved Contrastive Audio-Text Model
- Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- Zipformer: A faster and better encoder for automatic speech recognition
- Efficient Supervised Training of Audio Transformers for Music Representation Learning
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
- Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning
- EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
- SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality
- MARBLE: Music Audio Representation Benchmark for Universal Evaluation
- Image Captioners Are Scalable Vision Learners Too
- MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
- Listen, Think, and Understand
- DataComp: In search of the next generation of multimodal datasets
- Visual Instruction Tuning
- BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic Data
- MuChoMusic dataset
- BEATs: Audio Pre-Training with Acoustic Tokenizers
- Robust Speech Recognition via Large-Scale Weak Supervision
- Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
- LAION-5B: An open large-scale dataset for training next generation image-text models
- When and why vision-language models behave like bags-of-words, and what to do about it?
- Contrastive Audio-Visual Masked Autoencoder
- Masked Autoencoders that Listen
- What's in a Caption? Dataset-Specific Linguistic Diversity and Its Effect on Visual Description Models and Metrics
- Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
- VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance
- HEAR: Holistic Evaluation of Audio Representations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection
- Threshold Independent Evaluation of Sound Event Detection Scores
- Towards Learning Universal Audio Representations
- LiT: Zero-Shot Transfer with Locked-image text Tuning
- Lhotse: a speech data representation library for the modern deep learning ecosystem
- Wav2CLIP: Learning Robust Audio Representations From CLIP
- SSAST: Self-Supervised Audio Spectrogram Transformer
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- The Benefit Of Temporally-Strong Labels In Audio Event Classification
- SUPERB: Speech processing Universal PERformance Benchmark
- AST: Audio Spectrogram Transformer
- Learning Transferable Visual Models From Natural Language Supervision
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize\n Long-Tail Visual Concepts
- FSD50K: An Open Dataset of Human-Labeled Sound Events
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- VGGSound: A Large-scale Audio-Visual Dataset
- Voxceleb: Large-scale speaker verification in the wild
- A Simple Framework for Contrastive Learning of Visual Representations
- PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
- Common Voice: A Massively-Multilingual Speech Corpus
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language\n Generation, Translation, and Comprehension
- Clotho: An Audio Captioning Dataset
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in\n Conversations
- Representation Learning with Contrastive Predictive Coding
- Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
- CNN Architectures for Large-Scale Audio Classification
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Pointer Sentinel Mixture Models
- A Diversity-Promoting Objective Function for Neural Conversation Models
- Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related