vix.ing · top · new · best · stats · spec

Variable frame rate-based data augmentation to handle speaking-style\n variability for automatic speaker verification

2020/08/08 by Amber Afshan, Jinxi Guo, Afshan, Amber +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2008.03616

openalex publication_date 2020/08/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The effects of speaking-style variability on automatic speaker verification\nwere investigated using the UCLA Speaker Variability database which comprises\nmultiple speaking styles per speaker. An x-vector/PLDA (probabilistic linear\ndiscriminant analysis) system was trained with the SRE and Switchboard\ndatabases with standard augmentation techniques and evaluated with utterances\nfrom the UCLA database. The equal error rate (EER) was low when enrollment and\ntest utterances were of the same style (e.g., 0.98% and 0.57% for read and\nconversational speech, respectively), but it increased substantially when\nstyles were mismatched between enrollment and test utterances. For instance,\nwhen enrolled with conversation utterances, the EER increased to 3.03%, 2.96%\nand 22.12% when tested on read, narrative, and pet-directed speech,\nrespectively. To reduce the effect of style mismatch, we propose an\nentropy-based variable frame rate technique to artificially generate\nstyle-normalized representations for PLDA adaptation. The proposed system\nsignificantly improved performance. In the aforementioned conditions, the EERs\nimproved to 2.69% (conversation -- read), 2.27% (conversation -- narrative),\nand 18.75% (pet-directed -- read). Overall, the proposed technique performed\ncomparably to multi-style PLDA adaptation without the need for training data in\ndifferent speaking styles per speaker.\n

Cited by

Related