Advances in Speech Separation: Techniques, Challenges, and Future Trends
2025/08/14 by Li, Kai, Chen, Guo, Sang, Wendi +8 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2508.10830
Abstract
The field of speech separation, addressing the "cocktail party problem", has seen revolutionary advances with DNNs. Speech separation enhances clarity in complex acoustic environments and serves as crucial pre-processing for speech recognition and speaker recognition. However, current literature focuses narrowly on specific architectures or isolated approaches, creating fragmented understanding. This survey addresses this gap by providing systematic examination of DNN-based speech separation techniques. Our work differentiates itself through: (I) Comprehensive perspective: We systematically investigate learning paradigms, separation scenarios with known/unknown speakers, comparative analysis of supervised/self-supervised/unsupervised frameworks, and architectural components from encoders to estimation strategies. (II) Timeliness: Coverage of cutting-edge developments ensures access to current innovations and benchmarks. (III) Unique insights: Beyond summarization, we evaluate technological trajectories, identify emerging patterns, and highlight promising directions including domain-robust frameworks, efficient architectures, multimodal integration, and novel self-supervised paradigms. (IV) Fair evaluation: We provide quantitative evaluations on standard datasets, revealing true capabilities and limitations of different methods. This comprehensive survey serves as an accessible reference for experienced researchers and newcomers navigating speech separation's complex landscape.
Citations
- ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
- Listen to Extract: Onset-Prompted Target Speaker Extraction
- EDSep: An Effective Diffusion-Based Method for Speech Source Separation
- Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
- SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
- Speech Separation using Neural Audio Codecs with Embedding Loss
- Multi-Level Speaker Representation for Target Speaker Extraction
- TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
- SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
- Target Speaker ASR with Whisper
- WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
- LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
- Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
- Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
- Open-Source Conversational AI with SpeechBrain 1.0
- Towards Audio Codec-based Speech Separation
- Noise-robust Speech Separation with Fast Generative Correction
- ICASSP 2024 Speech Signal Improvement Challenge
- Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor
- MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
- USED: Universal Speaker Extraction and Diarization
- Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
- ReZero: Region-customizable Sound Extraction
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
- SpatialNet: Extensively Learning Spatial Information for Multichannel Joint Speech Separation, Denoising and Dereverberation
- Mixture Encoder for Joint Speech Separation and Recognition
- UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training Mixtures
- Cocktail HuBERT: Generalized Self-Supervised Pre-training for Mixture and Single-Source Speech
- MossFormer: Pushing the Performance Limit of Monaural Speech Separation using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions
- Neural Target Speech Extraction: An overview
- Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation
- An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits
- Deep neural network techniques for monaural speech enhancement: state of the art analysis
- TF-GridNet: Integrating Full- and Sub-Band Modeling for Speech Separation
- Diffusion-based Generative Speech Source Separation
- Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit
- Hierarchical speaker representation for target speaker extraction
- High Fidelity Neural Audio Compression
- Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features
- An efficient encoder-decoder architecture with top-down attention for speech separation
- Music Source Separation with Band-split RNN
- VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
- FRA-RIR: Fast Random Approximation of the Image-source Method
- ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding
- Resource-Efficient Separation Transformer
- HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
- EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers
- Investigating self-supervised learning for speech enhancement and separation
- L-SpEx: Localized Target Speaker Extraction
- SkiM: Skipping Memory LSTM for Low-Latency Real-Time Continuous Speech Separation
- Class-aware Sounding Objects Localization via Audiovisual Correspondence
- Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural Network
- Mixed Precision DNN Qunatization for Overlapped Speech Separation and Recognition
- Weight, Block or Unit? Exploring Sparsity Tradeoffs for Speech Enhancement on Tiny Neural Accelerators
- SA-SDR: A novel loss function for separation of meeting style data
- REAL-M: Towards Speech Separation on Real Mixtures
- M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
- Generative Adversarial Networks
- SoundStream: An End-to-End Neural Audio Codec
- Investigation of Practical Aspects of Single Channel Speech Separation for ASR
- Teacher-Student MixIT for Unsupervised and Semi-supervised Speech Separation
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- SpeechBrain: A General-Purpose Speech Toolkit
- RegionViT: Regional-to-Local Attention for Vision Transformers
- AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,\n Recognition and Speaker Diarization in Conference Scenario
- Sandglasset: A Light Multi-Granularity Self-attentive Network For Time-Domain Speech Separation
- Medical Transformer: Gated Axial-Attention for Medical Image Segmentation
- A Review of Speaker Diarization: Recent Advances with Deep Learning
- A review of speaker diarization: Recent advances with deep learning
- Interspeech 2021 Deep Noise Suppression Challenge
- Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals
- Single channel voice separation for unknown number of speakers under\n reverberant and noisy settings
- Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
- Attention is All You Need in Speech Separation
- Muse: Multi-modal target speaker extraction with visual cues
- An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and\n Separation
- Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- Denoising Diffusion Probabilistic Models
- Conv-Linformer: Boosting Linformer's Performance with Convolution in Small-Scale Settings
- LibriMix: An Open-Source Dataset for Generalizable Speech Separation
- The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results
- Asteroid: the PyTorch-based audio source separation toolkit for\n researchers
- Voice Separation with an Unknown Number of Multiple Speakers
- Wavesplit: End-to-End Speech Separation by Speaker Clustering
- On Cross-Corpus Generalization of Deep Learning Based Speech Enhancement
- Continuous speech separation: dataset and analysis
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Onssen: an open-source speech separation and enhancement library
- SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition
- Analyzing the impact of speaker localization errors on speech separation for automatic speech recognition
- Filterbank design for end-to-end speech separation
- WHAMR!: Noisy and Reverberant Single-Channel Speech Separation
- Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation
- FaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing
- Axial Attention in Multidimensional Transformers
- WHAM!: Extending Speech Separation to Noisy Environments
- A Review of Recurrent Neural Networks: LSTM Cells and Network Architectures
- MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement
- Divide and Conquer: A Deep CASA Approach to Talker-independent Monaural Speaker Separation
- Recursive speech separation for unknown number of speakers
- FurcaNeXt: End-to-end monaural speech separation with dynamic gated dilated temporal convolutional networks
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- Efficient Attention: Attention with Linear Complexities
- Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective
- SDR - half-baked or well done?
- VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
- MGGAN: Solving Mode Collapse using Manifold Guided Training
- TasNet: time-domain audio separation network for real-time, single-channel speech separation
- SVSGAN: Singing Voice Separation via Generative Adversarial Network
- Generative Adversarial Source Separation
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Simple Recurrent Units for Highly Parallelizable Recurrence
- Supervised Speech Separation Based on Deep Learning: An Overview
- Attention Is All You Need
- Multi-talker Speech Separation with Utterance-level Permutation\n Invariant Training of Deep Recurrent Neural Networks
- Image De-raining Using a Conditional Generative Adversarial Network
- Least Squares Generative Adversarial Networks
- Temporal Convolutional Networks: A Unified Approach to Action\n Segmentation
- Single-Channel Multi-Speaker Separation using Deep Clustering
- Permutation Invariant Training of Deep Models for Speaker-Independent Multi-talker Speech Separation
- Deep clustering: Discriminative embeddings for segmentation and separation
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Independent component analysis: algorithms and applications
- SPMamba: State-space model is all you need in speech separation
- Diffusion-based Signal Refiner for Speech Separation
Cited by
Related