2021/01/01 by Sung Hwan Mun, Min Hyun Han, Dongjune Lee +3 · 5 citations
Computer Science · Engineering · #Artificial intelligence #Computer science #Embedding #Feature learning #Machine learning #Music and Audio Processing #Pattern recognition (psychology) #Probabilistic logic #Regularization (linguistics) #Representation (politics) #Speaker diarisation #Speaker recognition #Speech Recognition and Synthesis #Speech and Audio Processing #Speech recognition #cs.AI #cs.LG #cs.SD #eess.AS
paper · pdf · doi:10.1109/access.2021.3137190
published in IEEE Access 9, 167615-167627 (Institute of Electrical and Electronics Engineers) · Accepted by IEEE Access
openalex publication_date 2021/01/01 · arxiv created 2021/12/24 · arxiv updated 2021/12/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and an uncertainty-aware probabilistic speaker embedding training in the back-end. In the front-end stage, we learn the speaker representations via the bootstrap training scheme with the uniformity regularization term. In the back-end stage, the probabilistic speaker embeddings are estimated by maximizing the mutual likelihood score between the speech samples belonging to the same speaker, which provide not only speaker representations but also data uncertainty. Experimental results show that the proposed bootstrap equilibrium training strategy can effectively help learn the speaker representations and outperforms the conventional methods based on contrastive learning. Also, we demonstrate that the integrated two-stage framework further improves the speaker verification performance on the VoxCeleb1 test set in terms of EER and MinDCF.