JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
2025/12/14 by Chao, Jianghan, Gao, Jianzhang, Tan, Wenhui +3
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
paper · doi:10.48550/arxiv.2512.12772
Abstract
Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio, an effective benchmark must comprehensively cover three key aspects: (1) multi-modal dependency (i.e., questions that cannot be answered using vision or audio alone), (2) diverse audio information types (e.g., speech, sound events), and (3) varying scene spans. However, existing datasets fall short in one or more of these dimensions, limiting strict and comprehensive evaluation. To address this gap, we introduce JointAVBench, a novel benchmark with strict audio-video correlation, spanning five cognitive dimensions, four audio information types (speech, sound events, music, vocal traits), and three scene spans (single-, cross-, and full-scene). Given the high cost of manual annotation, we propose an automated pipeline that leverages state-of-the-art vision-LLMs, audio-LLMs, and general-purpose LLMs to synthesize questions and answers that strictly require joint audio-visual understanding. We evaluate leading vision-only, audio-only, and Omni-LLMs on our dataset. Results show that even the best-performing Omni-LLM achieves an average accuracy of only 62.6%, outperforming uni-modal baselines but revealing substantial room for improvement, especially in cross-scene reasoning.
Citations
- Qwen3-Omni Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Kimi-Audio Technical Report
- Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
- Qwen2.5-Omni Technical Report
- Qwen2.5-VL Technical Report
- video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- Qwen2.5 Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
- StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification
- GPT-4o System Card
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- VITA: Towards Open-Source Interactive Omni Multimodal LLM
- Qwen2-Audio Technical Report
- Tarsier: Recipes for Training and Evaluating Large Video Description Models
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
- MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
- LVBench: An Extreme Long Video Understanding Benchmark
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
- CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
- TempCompass: Do Video LLMs Really Understand Videos?
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
- SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- Audio-Visual LLM for Video Understanding
- OneLLM: One Framework to Align All Modalities with Language
- VTimeLLM: Empower LLM to Grasp Video Moments
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- NExT-GPT: Any-to-Any Multimodal LLM
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- PandaGPT: One Model To Instruction-Follow Them All
- GPT-4 Technical Report
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Robust Speech Recognition via Large-Scale Weak Supervision
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- Learning to Answer Questions in Dynamic Audio-Visual Scenarios
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Learning Transferable Visual Models From Natural Language Supervision
- TALL: Temporal Activity Localization via Language Query
- The Measurement of Observer Agreement for Categorical Data
- Audio-centric Video Understanding Benchmark without Text Shortcut
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- Long Story Short: Story-level Video Understanding from 20K Short Films
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
Related