WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
2022/07/04 by Sanyuan Chen, Chengyi Wang, Zhengyang Chen +15 · 232 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
paper · doi:10.1109/jstsp.2022.3188113
openalex publication_date 2022/07/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
Self-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks. As speech signal contains multi-faceted information including speaker identity, paralinguistics, spoken content, etc., learning universal representations for all speech tasks is challenging. To tackle the problem, we propose a new pre-trained model, WavLM, to solve full-stack downstream speech tasks. WavLM jointly learns masked speech prediction and denoising in pre-training. By this means, WavLM does not only keep the speech content modeling capability by the masked speech prediction, but also improves the potential to non-ASR tasks by the speech denoising. In addition, WavLM employs gated relative position bias for the Transformer structure to better capture the sequence ordering of input speech. We also scale up the training dataset from 60 k hours to 94 k hours. WavLM Large achieves state-of-the-art performance on the SUPERB benchmark, and brings significant improvements for various speech processing tasks on their representative benchmarks.
Cited by
- On feature representations for marmoset vocal communication analysis
- Cross-modal enhancement of speech representations via textual supervision for paralinguistic analysis
- Bayesian Speech synthesizers Can Learn from Multiple Teachers
- emg2speech: synthesizing speech from electromyography using self-supervised speech models
- Adapting Speech Foundation Models with Large Language Models for Unified Speech Recognition
- Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
- ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
- Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
- UniSE: A Unified Framework for Decoder-only Autoregressive LM-based Speech Enhancement
- Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Compositional domain adaptation for automatic speech recognition with headwise selective attention merging
- VBx for End-to-End Neural and Clustering-based Diarization
- DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
- SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
- SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
- RLAIF-SPA: Optimizing LLM-based Emotional Speech Synthesis via RLAIF
- Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
- Switchboard-Affect: Emotion Perception Labels from Conversational Speech
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
- Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker
- FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
- MelTok: 2D Tokenization for Single-Codebook Audio Compression
- Joint Optimization of Speaker and Spoof Detectors for Spoofing-Robust Automatic Speaker Verification
- Unsupervised lexicon learning from speech is limited by representations rather than clustering
- OO-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
- FLToP CTC: Frame-Level Token Pruning via Relative Threshold for Efficient and Memory-Saving Decoding on Diverse Platforms
- IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
- Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
- Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
- Position: Towards Responsible Evaluation for Text-to-Speech
- TokenChain: A Discrete Speech Chain via Semantic Token Modeling
- MuFFIN: Multifaceted Pronunciation Feedback Model with Interactive Hierarchical Neural Modeling
- Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
- Provable Speech Attributes Conversion via Latent Independence
- Social Agent: Mastering Dyadic Nonverbal Behavior Generation via Conversational LLM Agents
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- Evaluating Self-Supervised Speech Models via Text-Based LLMS
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
- Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
- Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
- Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
- Towards Unsupervised Speech Recognition at the Syllable-Level
- STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Backdoor Attacks Against Speech Language Models
- FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
- Reference-free automatic speech severity evaluation using acoustic unit language modelling
- MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
- Scaling Spoken Language Models with Syllabic Speech Tokenization
- On Deepfake Voice Detection -- It's All in the Presentation
- Optimizing Speech Language Models for Acoustic Consistency
- Benchmarking Diarization Models
- Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
- Sparse Autoencoders Make Audio Foundation Models more Explainable
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
- SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement
- Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
- SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
- AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
- MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
- A novel fusion architecture for detecting Parkinson’s Disease using semi-supervised speech embeddings
- WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
- AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
- Frustratingly Easy Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching
- HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
- ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
- TF-Restormer: Complex Spectral Prediction for Speech Restoration
- The Impact of Audio Watermarking on Audio Anti-Spoofing Countermeasures
- MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
- PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
- Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
- Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
- CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
- MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
- Short-Segment Speaker Verification with Pre-trained Models and Multi-Resolution Encoder
- Selective Classifier-free Guidance for Zero-shot Text-to-speech
- Teffic-Audio: Tell Fact from Fiction
- VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- Enhancing Speech Emotion Recognition with Multi-Task Learning and Dynamic Feature Fusion
- Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
- HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
- M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
- STAR: Speech-to-Audio Generation via Representation Learning
- MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
- EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
- Multi-task Pretraining for Enhancing Interpretable L2 Pronunciation Assessment
- SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions
- MBCodec:Thorough disentangle for high-fidelity audio compression
- MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
- Speaker Style-Aware Phoneme Anchoring for Improved Cross-Lingual Speech Emotion Recognition
- Whisper-UT: A Unified Translation Framework for Speech and Text
- FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
- Are Multimodal Foundation Models All That Is Needed for Emofake Detection?
- Rethinking Cross-Corpus Speech Emotion Recognition Benchmarking: Are Paralinguistic Pre-Trained Representations Sufficient?
- Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations
- EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
- Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
- LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
- Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
- Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
- BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
- Llama-Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens
- MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
- HARNESS: Lightweight Distilled Arabic Speech Foundation Models
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
- CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
- Discrete optimal transport is a strong audio adversarial attack
- A long-form single-speaker real-time MRI speech dataset and benchmark
- VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
- SpeechOp: Inference-Time Task Composition for Generative Speech Processing
- SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification
- TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
- Improving Out-of-Domain Audio Deepfake Detection via Layer Selection and Fusion of SSL-Based Countermeasures
- More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition
- In-domain SSL pre-training and streaming ASR
- SENSE models: an open source solution for multilingual and multimodal semantic-based tasks
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
- Length-Aware Rotary Position Embedding for Text-Speech Alignment
- An Entropy-Guided Curriculum Learning Strategy for Data-Efficient Acoustic Scene Classification under Domain Shift
- Error Analysis in a Modular Meeting Transcription System
- The MSP-Podcast Corpus
- DiFlow-TTS: Discrete Flow Matching with Factorized Speech Tokens for Low-Latency Zero-Shot Text-To-Speech
- Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal Insufficiency
- Deep Learning for Tuberculosis Screening in a High-burden Setting using Cough Analysis and Speech Foundation Models
- Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
- MAPSS: Manifold-based Assessment of Perceptual Source Separation
- DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
- MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
- Audio Deepfake Verification
- Joint Learning using Mixture-of-Expert-Based Representation for Enhanced Speech Generation and Robust Emotion Recognition
- A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
- The First Voice Timbre Attribute Detection Challenge
- Continuous Audio Language Models
- Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet
- Layer-wise Analysis for Quality of Multilingual Synthesized Speech
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
- Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation
- Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
- Wav2DF-TSL: Two-stage Learning with Efficient Pre-training and Hierarchical Experts Fusion for Robust Audio Deepfake Detection
- Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
- Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
- Multi-level SSL Feature Gating for Audio Deepfake Detection
- Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
- AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
- The AudioMOS Challenge 2025
- Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
- A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
- Entropy-based Coarse and Compressed Semantic Speech Representation Learning
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
- Generalizable Audio Spoofing Detection using Non-Semantic Representations
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
- Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
- Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
- Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
- Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System
- Speech Emotion Recognition via Entropy-Aware Score Selection
- VibeVoice Technical Report
- Interpolating Speaker Identities in Embedding Space for Data Expansion
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- An Introduction to Silent Paralinguistics
- EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
- Cryfish: On deep audio analysis with Large Language Models
- HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
- Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
- VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
- Representing Speech Through Autoregressive Prediction of Cochlear Tokens
- Emphasis Sensitivity in Speech Representations
- Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
- Speech Emotion Recognition Using Fine-Tuned DWFormer:A Study on Track 1 of the IERPChallenge 2024
- Benchmarking Prosody Encoding in Discrete Speech Tokens
- Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
- Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
- Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
- Towards Frame-level Quality Predictions of Synthetic Speech
- ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
- Flow-SLM: Joint Learning of Linguistic and Acoustic Information for Spoken Language Modeling
- Exploring contrastive alignment across conversational turns for modeling vocal entrainment in interactions involving children with autism
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
- Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
- Transient Noise Removal via Diffusion-based Speech Inpainting
- Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
- Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
- Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
- Scalable Controllable Accented TTS
- KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
- Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
- ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
- SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
- EchoFree: Towards Ultra Lightweight and Efficient Neural Acoustic Echo Cancellation
- DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
- A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
- REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
- Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS
- Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
- LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model
- SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
- Inference-time Scaling for Diffusion-based Audio Super-resolution
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation
- Large Language Model Guided Decoding for Self-Supervised Speech Recognition
- Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
- Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life
- CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
- From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
- Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
- Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)
- Continual Speaker Identity Unlearning with Minimal Interference
Related