2017/10/06 by Mirco Ravanelli, Ravanelli, Mirco, Maurizio Omologo +1 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1710.02560
openalex publication_date 2017/10/06 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
This paper introduces the contents and the possible usage of the\nDIRHA-ENGLISH multi-microphone corpus, recently realized under the EC DIRHA\nproject. The reference scenario is a domestic environment equipped with a large\nnumber of microphones and microphone arrays distributed in space.\n The corpus is composed of both real and simulated material, and it includes\n12 US and 12 UK English native speakers. Each speaker uttered different sets of\nphonetically-rich sentences, newspaper articles, conversational speech,\nkeywords, and commands. From this material, a large set of 1-minute sequences\nwas generated, which also includes typical domestic background noise as well as\ninter/intra-room reverberation effects. Dev and test sets were derived, which\nrepresent a very precious material for different studies on multi-microphone\nspeech processing and distant-speech recognition. Various tasks and\ncorresponding Kaldi recipes have already been developed.\n The paper reports a first set of baseline results obtained using different\ntechniques, including Deep Neural Networks (DNN), aligned with the\nstate-of-the-art at international level.\n