2020/02/05 by Luciana Ferrer, Ferrer, Luciana, Mitchell McLaren +1 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2002.03802
openalex publication_date 2020/02/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In a recent work, we presented a discriminative backend for speaker\nverification that achieved good out-of-the-box calibration performance on most\ntested conditions containing varying levels of mismatch to the training\nconditions. This backend mimics the standard PLDA-based backend process used in\nmost current speaker verification systems, including the calibration stage. All\nparameters of the backend are jointly trained to optimize the binary\ncross-entropy for the speaker verification task. Calibration robustness is\nachieved by making the parameters of the calibration stage a function of\nvectors representing the conditions of the signal, which are extracted using a\nmodel trained to predict condition labels. In this work, we propose a\nsimplified version of this backend where the vectors used to compute the\ncalibration parameters are estimated within the backend, without the need for a\ncondition prediction model. We show that this simplified method provides\nsimilar performance to the previously proposed method while being simpler to\nimplement, and having less requirements on the training data. Further, we\nprovide an analysis of different aspects of the method including the effect of\ninitialization, the nature of the vectors used to compute the calibration\nparameters, and the effect that the random seed and the number of training\nepochs has on performance. We also compare the proposed method with the\ntrial-based calibration (TBC) method that, to our knowledge, was the\nstate-of-the-art for achieving good calibration across varying conditions. We\nshow that the proposed method outperforms TBC while also being several orders\nof magnitude faster to run, comparable to the standard PLDA baseline.\n