Generic Speech Enhancement with Self-Supervised Representation Space Loss
2025/07/10 by Sato, Hiroshi, Ochiai, Tsubasa, Delcroix, Marc +3
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2507.07631
Abstract
Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task. Thus, generalizing speech enhancement models to unknown downstream tasks has been challenging. This study aims to construct a generic speech enhancement front-end that can improve the performance of back-ends to solve multiple downstream tasks. To this end, we propose a novel training criterion that minimizes the distance between the enhanced and the ground truth clean signal in the feature representation domain of self-supervised learning models. Since self-supervised learning feature representations effectively express high-level speech information useful for solving various downstream tasks, the proposal is expected to make speech enhancement models preserve such information. Experimental validation demonstrates that the proposal improves the performance of multiple speech tasks while maintaining the perceptual quality of the enhanced signal.
Citations
- SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
- MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
- Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
- Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
- Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
- Joint Prediction and Denoising for Large-scale Multilingual Self-supervised Learning
- Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss
- ICASSP 2023 Deep Noise Suppression Challenge
- End-to-End Speech Recognition: A Survey
- VoxSRC 2022: The Fourth VoxCeleb Speaker Recognition Challenge
- Robust Speech Recognition via Large-Scale Weak Supervision
- End-to-End Integration of Speech Recognition, Dereverberation, Beamforming, and Self-Supervised Learning Representation
- TF-GridNet: Making Time-Frequency Domain Models Great Again for Monaural Speaker Separation
- ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding
- FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
- Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation
- End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning Representation
- SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities
- ICASSP 2022 Deep Noise Suppression Challenge
- How Bad Are Artifacts?: Analyzing the Impact of Speech Enhancement Errors on ASR
- A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation,\n Speech Enhancement and Speech Separation
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
- DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT
- Layer-wise Analysis of a Self-supervised Speech Representation Model
- HuBERT: Self-Supervised Speech Representation Learning by Masked\n Prediction of Hidden Units
- SUPERB: Speech processing Universal PERformance Benchmark
- The Zero Resource Speech Challenge 2021: Spoken language modelling
- TSTNN: Two-stage Transformer based Neural Network for Speech Enhancement in the Time Domain
- Interspeech 2021 Deep Noise Suppression Challenge
- Dual Application of Speech Enhancement for Automatic Speech Recognition
- A Survey on Contrastive Self-supervised Learning
- VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
- DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement
- Real Time Speech Enhancement in the Waveform Domain
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- Self-supervised Learning: Generative or Contrastive
- The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets,\n Subjective Testing Framework, and Challenge Results
- Improving noise robust automatic speech recognition with single-channel time-domain enhancement network
- Weighted Speech Distortion Losses for Neural-network-based Real-time Speech Enhancement
- Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- VoiceID Loss: Speech Enhancement for Speaker Verification
- SDR - half-baked or well done?
- VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
- ESPnet: End-to-End Speech Processing Toolkit
- Building state-of-the-art distant speech recognition using the CHiME-4 challenge with a setup of speech enhancement baseline
- Adam: A Method for Stochastic Optimization
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related