vix.ing · top · new · best · stats · spec

SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition

2019/10/30 by Lukas Drude, Jens Heitkaemper, Drude, Lukas +5 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1910.13934

openalex publication_date 2019/10/30 · openalex created_date 2019/11/08 · openalex updated_date 2026/07/28

Abstract

We present a multi-channel database of overlapping speech for training, evaluation, and detailed analysis of source separation and extraction algorithms: SMS-WSJ -- Spatialized Multi-Speaker Wall Street Journal. It consists of artificially mixed speech taken from the WSJ database, but unlike earlier databases we consider all WSJ0+1 utterances and take care of strictly separating the speaker sets present in the training, validation and test sets. When spatializing the data we ensure a high degree of randomness w.r.t. room size, array center and rotation, as well as speaker position. Furthermore, this paper offers a critical assessment of recently proposed measures of source separation performance. Alongside the code to generate the database we provide a source separation baseline and a Kaldi recipe with competitive word error rates to provide common ground for evaluation.

Citations

Cited by

Related