vix.ing · top · new · best · stats · spec

WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

2024/09/24 by Shuai Wang, Wang, Shuai, Ke Zhang +15 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2409.15799

openalex publication_date 2024/09/24 · openalex created_date 2024/10/26 · openalex updated_date 2026/07/28

Abstract

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to its potential for various applications such as user-customized interfaces and hearing aids, or as a crutial front-end processing technologies for subsequential tasks such as speech recognition and speaker recongtion. However, there are currently few open-source toolkits or available pre-trained models for off-the-shelf usage. In this work, we introduce WeSep, a toolkit designed for research and practical applications in TSE. WeSep is featured with flexible target speaker modeling, scalable data management, effective on-the-fly data simulation, structured recipes and deployment support. The toolkit is publicly avaliable at \urlhttps://github.com/wenet-e2e/WeSep.

Cited by

Related