Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
2025/12/22 by Apoorv Vyas, Vyas, Apoorv, Heng-Jui Chang +21 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Sound (cs.SD) #Speech and Audio Processing
paper · doi:10.48550/arxiv.2512.19687
openalex publication_date 2025/12/22 · openalex created_date 2025/12/24 · openalex updated_date 2026/07/28
Abstract
We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend representations to audio, and natively support joint embeddings across audio-video, audio-text, and video-text modalities. PE-AV's unified cross-modal embeddings enable novel tasks such as speech retrieval, and set a new state of the art across standard audio and video benchmarks. We unlock this by building a strong audiovisual data engine that synthesizes high-quality captions for O(100M) audio-video pairs, enabling large-scale supervision consistent across modalities. Our audio data includes speech, music, and general sound effects-avoiding single-domain limitations common in prior work. We exploit ten pairwise contrastive objectives, showing that scaling cross-modality and caption-type pairs strengthens alignment and improves zero-shot performance. We further develop PE-A-Frame by fine-tuning PE-AV with frame-level contrastive objectives, enabling fine-grained audio-frame-to-text alignment for tasks such as sound event detection.
Citations
- SAM Audio: Segment Anything in Audio
- FlexSED: Towards Open-Vocabulary Sound Event Detection
- Qwen3-Omni Technical Report
- USAD: Universal Speech and Audio Representation via Distillation
- FLAM: Frame-Wise Language-Audio Modeling
- Perception Encoder: The best visual embeddings are not at the output of the network
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
- DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
- A Comparative Study of Clinical ModernBERT and BioMedical ModernBERT on the DDXPlus Dataset
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
- Altogether: Image Captioning via Re-aligning Alt-text
- Movie Gen: A Cast of Media Foundation Models
- PaliGemma: A versatile 3B VLM for transfer
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
- LocCa: Visual Pretraining with Location-aware Captioners
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- VideoPrism: A Foundational Visual Encoder for Video Understanding
- EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
- EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- FLAP: Fast Language-Audio Pre-training
- SILC: Improving Vision Language Pretraining with Self-Distillation
- VeCLIP: Improving CLIP Training via Visual-enriched Captions
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
- Data Filtering Networks
- Demystifying CLIP Data
- Fine-tune the pretrained ATST model for sound event detection
- CoNeTTE: An efficient Audio Captioning system leveraging multiple datasets with Task Embedding
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
- Improving Multimodal Datasets with Image Captioning
- CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a $10,000 Budget; An Extra $4,000 Unlocks 81.8% Accuracy
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- Image Captioners Are Scalable Vision Learners Too
- Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks
- MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
- Improving CLIP Training with Language Rewrites
- Scaling Speech Technology to 1,000+ Languages
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
- ImageBind: One Embedding Space To Bind Them All
- DataComp: In search of the next generation of multimodal datasets
- VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
- Visual Instruction Tuning
- Unmasked Teacher: Towards Training-Efficient Video Foundation Models
- Sigmoid Loss for Language Image Pre-Training
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- BEATs: Audio Pre-Training with Acoustic Tokenizers
- MAViL: Masked Audio-Video Learners
- Scaling Language-Image Pre-training via Masking
- Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Contrastive Audio-Visual Masked Autoencoder
- Masked Autoencoders that Listen
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- Frequency Dynamic Convolution: Frequency-Adaptive Pattern Recognition for Sound Event Detection
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- Threshold Independent Evaluation of Sound Event Detection Scores
- SLIP: Self-supervision meets Language-Image Pre-training
- SSAST: Self-Supervised Audio Spectrogram Transformer
- Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- The Benefit Of Temporally-Strong Labels In Audio Event Classification
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Learning Transferable Visual Models From Natural Language Supervision
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- Learning Visual Representations with Caption Annotations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- VirTex: Learning Visual Representations from Textual Annotations
- VGGSound: A Large-scale Audio-Visual Dataset
- Common Voice: A Massively-Multilingual Speech Corpus
- Clotho: An Audio Captioning Dataset
- A Framework for the Robust Evaluation of Sound Event Detection
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- A Short Note about Kinetics-600
- Localizing Moments in Video with Natural Language
- Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
- UCF-101: A dataset of 101 human actions classes from videos in the wild
Cited by
Related