2020/08/09 by Florian Kreyssig, Kreyssig, Florian L., Philip C. Woodland +1
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2008.03756
openalex publication_date 2020/08/09 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
In this paper, we propose a semi-supervised learning (SSL) technique for\ntraining deep neural networks (DNNs) to generate speaker-discriminative\nacoustic embeddings (speaker embeddings). Obtaining large amounts of speaker\nrecognition train-ing data can be difficult for desired target domains,\nespecially under privacy constraints. The proposed technique reduces\nrequirements for labelled data by leveraging unlabelled data. The technique is\na variant of virtual adversarial training (VAT) [1] in the form of a loss that\nis defined as the robustness of the speaker embedding against input\nperturbations, as measured by the cosine-distance. Thus, we term the technique\ncosine-distance virtual adversarial training (CD-VAT). In comparison to many\nexisting SSL techniques, the unlabelled data does not have to come from the\nsame set of classes (here speakers) as the labelled data. The effectiveness of\nCD-VAT is shown on the 2750+ hour VoxCeleb data set, where on a speaker\nverification task it achieves a reduction in equal error rate (EER) of 11.1%\nrelative to a purely supervised baseline. This is 32.5% of the improvement that\nwould be achieved from supervised training if the speaker labels for the\nunlabelled data were available.\n