vix.ing · top · new · best · stats · spec

Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System

2024/07/13 by Lingwei Meng, Jiawen Kang, Meng, Lingwei +11 · 5 citations
Computer Science · Environmental Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Educational Reforms and Innovations #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2407.09817

openalex publication_date 2024/07/13 · openalex created_date 2024/07/17 · openalex updated_date 2026/07/28

Abstract

Multi-talker speech recognition and target-talker speech recognition, both involve transcription in multi-talker contexts, remain significant challenges. However, existing methods rarely attempt to simultaneously address both tasks. In this study, we propose a pioneering approach to empower Whisper, which is a speech foundation model, to tackle joint multi-talker and target-talker speech recognition tasks. Specifically, (i) we freeze Whisper and plug a Sidecar separator into its encoder to separate mixed embedding for multiple talkers; (ii) a Target Talker Identifier is introduced to identify the embedding flow of the target talker on the fly, requiring only three-second enrollment speech as a cue; (iii) soft prompt tuning for decoder is explored for better task adaptation. Our method outperforms previous methods on two- and three-talker LibriMix and LibriSpeechMix datasets for both tasks, and delivers acceptable zero-shot performance on multi-talker ASR on AishellMix Mandarin dataset.

Cited by

Related