vix.ing · top · new · best · stats · spec

Hear "No Evil", See "Kenansville": Efficient and Transferable Black-Box\n Attacks on Speech Recognition and Voice Identification Systems

2019/10/11 by Hadi Abdullah, Muhammad Sajidur Rahman, Abdullah, Hadi +13 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1910.05262

openalex publication_date 2019/10/11 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28

Abstract

Automatic speech recognition and voice identification systems are being\ndeployed in a wide array of applications, from providing control mechanisms to\ndevices lacking traditional interfaces, to the automatic transcription of\nconversations and authentication of users. Many of these applications have\nsignificant security and privacy considerations. We develop attacks that force\nmistranscription and misidentification in state of the art systems, with\nminimal impact on human comprehension. Processing pipelines for modern systems\nare comprised of signal preprocessing and feature extraction steps, whose\noutput is fed to a machine-learned model. Prior work has focused on the models,\nusing white-box knowledge to tailor model-specific attacks. We focus on the\npipeline stages before the models, which (unlike the models) are quite similar\nacross systems. As such, our attacks are black-box and transferable, and\ndemonstrably achieve mistranscription and misidentification rates as high as\n100% by modifying only a few frames of audio. We perform a study via Amazon\nMechanical Turk demonstrating that there is no statistically significant\ndifference between human perception of regular and perturbed audio. Our\nfindings suggest that models may learn aspects of speech that are generally not\nperceived by human subjects, but that are crucial for model accuracy. We also\nfind that certain English language phonemes (in particular, vowels) are\nsignificantly more susceptible to our attack. We show that the attacks are\neffective when mounted over cellular networks, where signals are subject to\ndegradation due to transcoding, jitter, and packet loss.\n

Cited by

Related