vix.ing · top · new · best · stats · spec

Adapting End-to-End Neural Speaker Verification to New Languages and\n Recording Conditions with Adversarial Training

2018/11/07 by Gautam Bhattacharya, Jahangir Alam, Bhattacharya, Gautam +3
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1811.03055

openalex publication_date 2018/11/07 · openalex created_date 2022/08/02 · openalex updated_date 2026/07/28

Abstract

In this article we propose a novel approach for adapting speaker embeddings\nto new domains based on adversarial training of neural networks. We apply our\nembeddings to the task of text-independent speaker verification, a challenging,\nreal-world problem in biometric security. We further the development of\nend-to-end speaker embedding models by combing a novel 1-dimensional,\nself-attentive residual network, an angular margin loss function and\nadversarial training strategy. Our model is able to learn extremely compact,\n64-dimensional speaker embeddings that deliver competitive performance on a\nnumber of popular datasets using simple cosine distance scoring. One the\nNIST-SRE 2016 task we are able to beat a strong i-vector baseline, while on the\nSpeakers in the Wild task our model was able to outperform both i-vector and\nx-vector baselines, showing an absolute improvement of 2.19% over the latter.\nAdditionally, we show that the integration of adversarial training consistently\nleads to a significant improvement over an unadapted model.\n

Related