vix.ing · top · new · best · stats · spec

Noise-tolerant Audio-visual Online Person Verification using an\n Attention-based Neural Network Fusion

2018/11/26 by Suwon Shon, Tae-Hyun Oh, Shon, Suwon +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Speech and Audio Processing #Video Surveillance and Tracking Methods #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1811.10813

openalex publication_date 2018/11/26 · openalex created_date 2022/08/01 · openalex updated_date 2026/07/28

Abstract

In this paper, we present a multi-modal online person verification system\nusing both speech and visual signals. Inspired by neuroscientific findings on\nthe association of voice and face, we propose an attention-based end-to-end\nneural network that learns multi-sensory associations for the task of person\nverification. The attention mechanism in our proposed network learns to\nconditionally select a salient modality between speech and facial\nrepresentations that provides a balance between complementary inputs. By virtue\nof this capability, the network is robust to missing or corrupted data from\neither modality. In the VoxCeleb2 dataset, we show that our method performs\nfavorably against competing multi-modal methods. Even for extreme cases of\nlarge corruption or an entirely missing modality, our method demonstrates\nrobustness over other unimodal methods.\n

Citations

Cited by

Related