2020/05/19 by François Grondin, Grondin, Francois, Jean-Samuel Lauzon +5
Computer Science · Engineering · #Speech and Audio Processing #Acoustic Wave Phenomena Research #Advanced Adaptive Filtering Techniques
paper · pdf · doi:10.48550/arxiv.2005.09587
Distant speech processing is a challenging task, especially when dealing with\nthe cocktail party effect. Sound source separation is thus often required as a\npreprocessing step prior to speech recognition to improve the signal to\ndistortion ratio (SDR). Recently, a combination of beamforming and speech\nseparation networks have been proposed to improve the target source quality in\nthe direction of arrival of interest. However, with this type of approach, the\nneural network needs to be trained in advance for a specific microphone array\ngeometry, which limits versatility when adding/removing microphones, or\nchanging the shape of the array. The solution presented in this paper is to\ntrain a neural network on pairs of microphones with different spacing and\nacoustic environmental conditions, and then use this network to estimate a\ntime-frequency mask from all the pairs of microphones forming the array with an\narbitrary shape. Using this mask, the target and noise covariance matrices can\nbe estimated, and then used to perform generalized eigenvalue (GEV)\nbeamforming. Results show that the proposed approach improves the SDR from 4.78\ndB to 7.69 dB on average, for various microphone array geometries that\ncorrespond to commercially available hardware.\n