2023/07/28 by Zhang, Li, Zhao, Huan, Li, Yue +6 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2307.15400
This paper describes the FlySpeech speaker diarization system submitted to the second Multimodal Information Based Speech Processing~(MISP) Challenge held in ICASSP 2022. We develop an end-to-end audio-visual speaker diarization~(AVSD) system, which consists of a lip encoder, a speaker encoder, and an audio-visual decoder. Specifically, to mitigate the degradation of diarization performance caused by separate training, we jointly train the speaker encoder and the audio-visual decoder. In addition, we leverage the large-data pretrained speaker extractor to initialize the speaker encoder.