2024/09/11 by Hong-Hanh Nguyen-Le, Van‐Tuan Tran, Nguyen-Le, Hong-Hanh +5
Computer Science · #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Information and Cyber Security #Machine Learning (cs.LG) #Sound (cs.SD) #User Authentication and Security Systems #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2409.07390
openalex publication_date 2024/09/11 · openalex created_date 2024/10/22 · openalex updated_date 2026/07/28
The advancements in generative AI have enabled the improvement of audio synthesis models, including text-to-speech and voice conversion. This raises concerns about its potential misuse in social manipulation and political interference, as synthetic speech has become indistinguishable from natural human speech. Several speech-generation programs are utilized for malicious purposes, especially impersonating individuals through phone calls. Therefore, detecting fake audio is crucial to maintain social security and safeguard the integrity of information. Recent research has proposed a D-CAPTCHA system based on the challenge-response protocol to differentiate fake phone calls from real ones. In this work, we study the resilience of this system and introduce a more robust version, D-CAPTCHA++, to defend against fake calls. Specifically, we first expose the vulnerability of the D-CAPTCHA system under transferable imperceptible adversarial attack. Secondly, we mitigate such vulnerability by improving the robustness of the system by using adversarial training in D-CAPTCHA deepfake detectors and task classifiers.