vix.ing · top · new · best · stats

Audio-Visual Event Localization in Unconstrained Videos

2018/03/23 by Yapeng Tian, Jing Shi, Tian, Yapeng +7 · 62 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing #Speech and Audio Processing #Video Analysis and Summarization #cs.CV

paper · pdf · doi:10.48550/arxiv.1803.08842

23 pages, 7 figures

arxiv created 2018/03/23 · openalex publication_date 2018/03/23 · arxiv updated 2018/03/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio-Visual Event(AVE) dataset to systemically investigate three temporal localization tasks: supervised and weakly-supervised audio-visual event localization, and cross-modality localization. We develop an audio-guided visual attention mechanism to explore audio-visual correlations, propose a dual multimodal residual network (DMRN) to fuse information over the two modalities, and introduce an audio-visual distance learning network to handle the cross-modality localization. Our experiments support the following findings: joint modeling of auditory and visual modalities outperforms independent modeling, the learned attention can capture semantics of sounding objects, temporal alignment is important for audio-visual fusion, the proposed DMRN is effective in fusing audio-visual features, and strong correlations between the two modalities enable cross-modality localization.

Cited by

Related