2021/07/12 by Anirudh Sreeram, Nicholas Mehlman, Sreeram, Anirudh +7
Computer Science · Engineering · #Speech Recognition and Synthesis #Adversarial Robustness in Machine Learning #Geophysical Methods and Applications
paper · pdf · doi:10.48550/arxiv.2107.05222
In this paper we investigate speech denoising as a defense against\nadversarial attacks on automatic speech recognition (ASR) systems. Adversarial\nattacks attempt to force misclassification by adding small perturbations to the\noriginal speech signal. We propose to counteract this by employing a\nneural-network based denoiser as a pre-processor in the ASR pipeline. The\ndenoiser is independent of the downstream ASR model, and thus can be rapidly\ndeployed in existing systems. We found that training the denoisier using a\nperceptually motivated loss function resulted in increased adversarial\nrobustness without compromising ASR performance on benign samples. Our defense\nwas evaluated (as a part of the DARPA GARD program) on the 'Kenansville' attack\nstrategy across a range of attack strengths and speech samples. An average\nimprovement in Word Error Rate (WER) of about 7.7% was observed over the\nundefended model at 20 dB signal-to-noise-ratio (SNR) attack strength.\n