vix.ing · top · new · best · stats

Self-Supervised Speaker Verification Using Dynamic Loss-Gate and Label Correction

2022/08/03 by Bing Han, Zhengyang Chen, Han, Bing +3 · 2 citations
Computer Science · Engineering · #Artificial intelligence #Audio and Speech Processing (eess.AS) #Computer science #Degradation (telecommunications) #FOS: Computer and information sciences #FOS: Electrical engineering #Gaussian #Mixture model #Music and Audio Processing #Pattern recognition (psychology) #Sound (cs.SD) #Speaker recognition #Speaker verification #Speech Recognition and Synthesis #Speech and Audio Processing #Speech recognition #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2208.01928

published in arXiv (Cornell University) (Cornell University) · Accepted by Interspeech 2022

arxiv created 2022/08/03 · openalex publication_date 2022/08/03 · arxiv updated 2022/08/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

For self-supervised speaker verification, the quality of pseudo labels decides the upper bound of the system due to the massive unreliable labels. In this work, we propose dynamic loss-gate and label correction (DLG-LC) to alleviate the performance degradation caused by unreliable estimated labels. In DLG, we adopt Gaussian Mixture Model (GMM) to dynamically model the loss distribution and use the estimated GMM to distinguish the reliable and unreliable labels automatically. Besides, to better utilize the unreliable data instead of dropping them directly, we correct the unreliable label with model predictions. Moreover, we apply the negative-pairs-free DINO framework in our experiments for further improvement. Compared to the best-known speaker verification system with self-supervised learning, our proposed DLG-LC converges faster and achieves 11.45%, 18.35% and 15.16% relative improvement on Vox-O, Vox-E and Vox-H trials of Voxceleb1 evaluation dataset.

Cited by

Related