2018/11/07 by Gautam Bhattacharya, João Monteiro, Bhattacharya, Gautam +5
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1811.03063
openalex publication_date 2018/11/07 · openalex created_date 2022/08/02 · openalex updated_date 2026/07/28
This article presents a novel approach for learning domain-invariant speaker\nembeddings using Generative Adversarial Networks. The main idea is to confuse a\ndomain discriminator so that is can't tell if embeddings are from the source or\ntarget domains. We train several GAN variants using our proposed framework and\napply them to the speaker verification task. On the challenging NIST-SRE 2016\ndataset, we are able to match the performance of a strong baseline x-vector\nsystem. In contrast to the the baseline systems which are dependent on\ndimensionality reduction (LDA) and an external classifier (PLDA), our proposed\nspeaker embeddings can be scored using simple cosine distance. This is achieved\nby optimizing our models end-to-end, using an angular margin loss function.\nFurthermore, we are able to significantly boost verification performance by\naveraging our different GAN models at the score level, achieving a relative\nimprovement of 7.2% over the baseline.\n