2022/04/18 by Chenyu Yang, Yu Wang, Yang, Chenyu +1 · 1 citation
Computer Science · Engineering · #Artificial intelligence #Artificial neural network #Cluster analysis #Computer science #End-to-end principle #Inference #Leverage (statistics) #Machine learning #Music and Audio Processing #Pattern recognition (psychology) #Speaker diarisation #Speaker recognition #Speech Recognition and Synthesis #Speech and Audio Processing #Speech recognition #cs.SD #eess.AS
paper · pdf · doi:10.48550/arxiv.2204.08164
published in arXiv (Cornell University) (Cornell University) · submitted to INTERSPEECH 2022
arxiv created 2022/04/18 · openalex publication_date 2022/04/18 · arxiv updated 2022/04/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
End-to-end speaker diarization approaches have shown exceptional performance over the traditional modular approaches. To further improve the performance of the end-to-end speaker diarization for real speech recordings, recently works have been proposed which integrate unsupervised clustering algorithms with the end-to-end neural diarization models. However, these methods have a number of drawbacks: 1) The unsupervised clustering algorithms cannot leverage the supervision from the available datasets; 2) The K-means-based unsupervised algorithms that are explored often suffer from the constraint violation problem; 3) There is unavoidable mismatch between the supervised training and the unsupervised inference. In this paper, a robust generic neural clustering approach is proposed that can be integrated with any chunk-level predictor to accomplish a fully supervised end-to-end speaker diarization model. Also, by leveraging the sequence modelling ability of a recurrent neural network, the proposed neural clustering approach can dynamically estimate the number of speakers during inference. Experimental show that when integrating an attractor-based chunk-level predictor, the proposed neural clustering approach can yield better Diarization Error Rate (DER) than the constrained K-means-based clustering approaches under the mismatched conditions.