vix.ing · top · new · best · stats · spec

Enhancement of Spatial Clustering-Based Time-Frequency Masks using LSTM\n Neural Networks

2020/12/02 by Félix Grèzes, Grezes, Felix, Zhaoheng Ni +5
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing

paper · pdf · doi:10.48550/arxiv.2012.01576

Abstract

Recent works have shown that Deep Recurrent Neural Networks using the LSTM\narchitecture can achieve strong single-channel speech enhancement by estimating\ntime-frequency masks. However, these models do not naturally generalize to\nmulti-channel inputs from varying microphone configurations. In contrast,\nspatial clustering techniques can achieve such generalization but lack a strong\nsignal model. Our work proposes a combination of the two approaches. By using\nLSTMs to enhance spatial clustering based time-frequency masks, we achieve both\nthe signal modeling performance of multiple single-channel LSTM-DNN speech\nenhancers and the signal separation performance and generality of multi-channel\nspatial clustering. We compare our proposed system to several baselines on the\nCHiME-3 dataset. We evaluate the quality of the audio from each system using\nSDR from the BSS\_eval toolkit and PESQ. We evaluate the intelligibility of the\noutput of each system using word error rate from a Kaldi automatic speech\nrecognizer.\n

Related