vix.ing · top · new · best · stats · spec

Combining Spatial Clustering with LSTM Speech Models for Multichannel\n Speech Enhancement

2020/12/02 by Félix Grèzes, Grezes, Felix, Zhaoheng Ni +5
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2012.03388

openalex publication_date 2020/12/02 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Recurrent neural networks using the LSTM architecture can achieve significant\nsingle-channel noise reduction. It is not obvious, however, how to apply them\nto multi-channel inputs in a way that can generalize to new microphone\nconfigurations. In contrast, spatial clustering techniques can achieve such\ngeneralization, but lack a strong signal model. This paper combines the two\napproaches to attain both the spatial separation performance and generality of\nmultichannel spatial clustering and the signal modeling performance of multiple\nparallel single-channel LSTM speech enhancers. The system is compared to\nseveral baselines on the CHiME3 dataset in terms of speech quality predicted by\nthe PESQ algorithm and word error rate of a recognizer trained on mis-matched\nconditions, in order to focus on generalization. Our experiments show that by\ncombining the LSTM models with the spatial clustering, we reduce word error\nrate by 4.6 % absolute (17.2 % relative) on the development set and 11.2 %\nabsolute (25.5 % relative) on test set compared with spatial clustering system,\nand reduce by 10.75 % (32.72 % relative) on development set and 6.12 % absolute\n(15.76 % relative) on test data compared with LSTM model.\n

Citations

Related