vix.ing · top · new · best · stats · spec

Deep Generative Variational Autoencoding for Replay Spoof Detection in\n Automatic Speaker Verification

2020/03/20 by Bhusan Chettri, Chettri, Bhusan, Tomi Kinnunen +3
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2003.09542

openalex publication_date 2020/03/20 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

Automatic speaker verification (ASV) systems are highly vulnerable to\npresentation attacks, also called spoofing attacks. Replay is among the\nsimplest attacks to mount - yet difficult to detect reliably. The\ngeneralization failure of spoofing countermeasures (CMs) has driven the\ncommunity to study various alternative deep learning CMs. The majority of them\nare supervised approaches that learn a human-spoof discriminator. In this\npaper, we advocate a different, deep generative approach that leverages from\npowerful unsupervised manifold learning in classification. The potential\nbenefits include the possibility to sample new data, and to obtain insights to\nthe latent features of genuine and spoofed speech. To this end, we propose to\nuse variational autoencoders (VAEs) as an alternative backend for replay attack\ndetection, via three alternative models that differ in their\nclass-conditioning. The first one, similar to the use of Gaussian mixture\nmodels (GMMs) in spoof detection, is to train independently two VAEs - one for\neach class. The second one is to train a single conditional model (C-VAE) by\ninjecting a one-hot class label vector to the encoder and decoder networks. Our\nfinal proposal integrates an auxiliary classifier to guide the learning of the\nlatent space. Our experimental results using constant-Q cepstral coefficient\n(CQCC) features on the ASVspoof 2017 and 2019 physical access subtask datasets\nindicate that the C-VAE offers substantial improvement in comparison to\ntraining two separate VAEs for each class. On the 2019 dataset, the C-VAE\noutperforms the VAE and the baseline GMM by an absolute 9 - 10% in both equal\nerror rate (EER) and tandem detection cost function (t-DCF) metrics. Finally,\nwe propose VAE residuals - the absolute difference of the original input and\nthe reconstruction as features for spoofing detection.\n

Citations

Related