vix.ing · top · new · best · stats · spec

Synchformer: Efficient Synchronization from Sparse Cues

2024/01/29 by Vladimir Iashin, Iashin, Vladimir, Weidi Xie +5 · 36 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Neural Networks and Applications #Neural Networks and Reservoir Computing #Photonic and Optical Devices #Sound (cs.SD) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2401.16423

openalex publication_date 2024/01/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that decouples feature extraction from synchronization modelling through multi-modal segment-level contrastive pre-training. This approach achieves state-of-the-art performance in both dense and sparse settings. We also extend synchronization model training to AudioSet a million-scale 'in-the-wild' dataset, investigate evidence attribution techniques for interpretability, and explore a new capability for synchronization models: audio-visual synchronizability.

Cited by

Related