vix.ing · top · new · best · stats · spec

Speaker Embedding Extraction with Phonetic Information

2018/04/13 by Yi Liu, Liu, Yi, Liang He +5 · 5 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1804.04862

submitted to Interspeech 2018 (accepted) and open-sourced. Please refer to Interspeech for the final version

openalex publication_date 2018/04/13 · arxiv created 2018/06/14 · arxiv updated 2018/06/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Speaker embeddings achieve promising results on many speaker verification tasks. Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings. In this paper, we introduce phonetic information to the speaker embedding extraction based on the x-vector architecture. Two methods using phonetic vectors and multi-task learning are proposed. On the Fisher dataset, our best system outperforms the original x-vector approach by 20% in EER, and by 15%, 15% in minDCF08 and minDCF10, respectively. Experiments conducted on NIST SRE10 further demonstrate the effectiveness of the proposed methods.

Citations

Cited by

Related