vix.ing · top · new · best · stats

Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation

2024/06/10 by Juhyeong Seon, Woobin Im, Seon, Juhyeong +7 · 3 citations
Computer Science · Neuroscience · Psychology · #Artificial intelligence #Audio visual #Communication #Computer Vision and Pattern Recognition (cs.CV) #Computer science #FOS: Computer and information sciences #Hearing Loss and Rehabilitation #Multimedia #Music and Audio Processing #Psychology #Segmentation #Speech and Audio Processing #Speech recognition

paper · pdf · doi:10.48550/arxiv.2406.06163

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/06/10 · openalex created_date 2024/06/12 · openalex updated_date 2026/07/28

Abstract

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense prediction problems, prior works have investigated the introduction of SAM into AVS with audio as a new modality of the prompt. Nevertheless, constrained by SAM's single-frame segmentation scheme, the temporal context across multiple frames of audio-visual data remains insufficiently utilized. To this end, we study the extension of SAM's capabilities to the sequence of audio-visual scenes by analyzing contextual cross-modal relationships across the frames. To achieve this, we propose a Spatio-Temporal, Bidirectional Audio-Visual Attention (ST-BAVA) module integrated into the middle of SAM's image encoder and mask decoder. It adaptively updates the audio-visual features to convey the spatio-temporal correspondence between the video frames and audio streams. Extensive experiments demonstrate that our proposed model outperforms the state-of-the-art methods on AVS benchmarks, especially with an 8.3% mIoU gain on a challenging multi-sources subset.

Cited by

Related