2021/03/31 by Anton Ratnarajah, Ratnarajah, Anton, Zhenyu Tang +3 · 2 citations
Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
paper · pdf · doi:10.48550/arxiv.2103.16804
We present a method for improving the quality of synthetic room impulse\nresponses for far-field speech recognition. We bridge the gap between the\nfidelity of synthetic room impulse responses (RIRs) and the real room impulse\nresponses using our novel, TS-RIRGAN architecture. Given a synthetic RIR in the\nform of raw audio, we use TS-RIRGAN to translate it into a real RIR. We also\nperform real-world sub-band room equalization on the translated synthetic RIR.\nOur overall approach improves the quality of synthetic RIRs by compensating\nlow-frequency wave effects, similar to those in real RIRs. We evaluate the\nperformance of improved synthetic RIRs on a far-field speech dataset augmented\nby convolving the LibriSpeech clean speech dataset [1] with RIRs and adding\nbackground noise. We show that far-field speech augmented using our improved\nsynthetic RIRs reduces the word error rate by up to 19.9% in Kaldi far-field\nautomatic speech recognition benchmark [2].\n