2018/11/26 by Suwon Shon, Tae-Hyun Oh, Shon, Suwon +3 · 1 citation
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Speech and Audio Processing #Video Surveillance and Tracking Methods #cs.CV #eess.AS #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1811.10813
openalex publication_date 2018/11/26 · arxiv created 2018/11/27 · arxiv updated 2018/11/28 · openalex created_date 2022/08/01 · openalex updated_date 2026/07/28
In this paper, we present a multi-modal online person verification system using both speech and visual signals. Inspired by neuroscientific findings on the association of voice and face, we propose an attention-based end-to-end neural network that learns multi-sensory associations for the task of person verification. The attention mechanism in our proposed network learns to conditionally select a salient modality between speech and facial representations that provides a balance between complementary inputs. By virtue of this capability, the network is robust to missing or corrupted data from either modality. In the VoxCeleb2 dataset, we show that our method performs favorably against competing multi-modal methods. Even for extreme cases of large corruption or an entirely missing modality, our method demonstrates robustness over other unimodal methods.