2020/02/09 by Raghuveer Peri, Haoqi Li, Peri, Raghuveer +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2002.03520
openalex publication_date 2020/02/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The primary characteristic of robust speaker representations is that they are\ninvariant to factors of variability not related to speaker identity.\nDisentanglement of speaker representations is one of the techniques used to\nimprove robustness of speaker representations to both intrinsic factors that\nare acquired during speech production (e.g., emotion, lexical content) and\nextrinsic factors that are acquired during signal capture (e.g., channel,\nnoise). Disentanglement in neural speaker representations can be achieved\neither in a supervised fashion with annotations of the nuisance factors\n(factors not related to speaker identity) or in an unsupervised fashion without\nlabels of the factors to be removed. In either case it is important to\nunderstand the extent to which the various factors of variability are entangled\nin the representations. In this work, we examine speaker representations with\nand without unsupervised disentanglement for the amount of information they\ncapture related to a suite of factors. Using classification experiments we\nprovide empirical evidence that disentanglement reduces the information with\nrespect to nuisance factors from speaker representations, while retaining\nspeaker information. This is further validated by speaker verification\nexperiments on the VOiCES corpus in several challenging acoustic conditions. We\nalso show improved robustness in speaker verification tasks using data\naugmentation during training of disentangled speaker embeddings. Finally, based\non our findings, we provide insights into the factors that can be effectively\nseparated using the unsupervised disentanglement technique and discuss\npotential future directions.\n