vix.ing · top · new · best · stats · spec

Simultaneous Speech Extraction for Multiple Target Speakers under the Meeting Scenarios

2022/06/17 by Bang Zeng, Zeng, Bang, Suo, Hongbing +3
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing

paper · pdf · doi:10.48550/arxiv.2206.08525

Abstract

The common target speech separation directly estimate the target source, ignoring the interrelationship between different speakers at each frame. We propose a multiple-target speech separation model (MTSS) to simultaneously extract each speaker's voice from the mixed speech rather than just optimally estimating the target source. Moreover, we propose a speaker diarization (SD) aware MTSS system (SD-MTSS), which consists of a SD module and MTSS module. By exploiting the TSVAD decision and the estimated mask, our SD-MTSS model can extract the speech signal of each speaker concurrently in a conversational recording without additional enrollment audio in advance. Experimental results show that our MTSS model achieves 1.38dB SDR, 1.34dB SI-SDR, and 0.13 PESQ improvements over the baseline on the WSJ0-2mix-extr dataset, respectively. The SD-MTSS system makes 19.2% relative speaker dependent character error rate (CER) reduction on the Alimeeting dataset.

Related