LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
2025/10/20 by Han, ZhaoYang, Lin, Qihan, Liang, Hao +3
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
paper · doi:10.48550/arxiv.2510.17305
Abstract
We introduce LongInsightBench, the first benchmark designed to assess models' ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements, while integrating visual, audio, and text modalities. Our benchmark excels in three key areas: a) Long-Duration, Information-Dense Videos: We carefully select approximately 1,000 videos from open-source datasets FineVideo based on duration limit and the information density of both visual and audio modalities, focusing on content like lectures, interviews, and vlogs, which contain rich language elements. b) Diverse and Challenging Task Scenarios: We have designed six challenging task scenarios, including both Intra-Event and Inter-Event Tasks. c) Rigorous and Comprehensive Quality Assurance Pipelines: We have developed a three-step, semi-automated data quality assurance pipeline to ensure the difficulty and validity of the synthesized questions and answer options. Based on LongInsightBench, we designed a series of experiments. Experimental results shows that Omni-modal models(OLMs) still face challenge in tasks requiring precise temporal localization (T-Loc) and long-range causal inference (CE-Caus). Extended experiments reveal the information loss and processing bias in multi-modal fusion of OLMs. Our dataset and code is available at https://anonymous.4open.science/r/LongInsightBench-910F/.
Citations
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
- Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
- Ola: Pushing the Frontiers of Omni-Modal Language Model
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- Towards Long Video Understanding via Fine-detailed Video Story Generation
- MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
- MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- ImageBind: One Embedding Space To Bind Them All
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
- Robust Speech Recognition via Large-Scale Weak Supervision
- Learning to Answer Questions in Dynamic Audio-Visual Scenarios
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- Learning Transferable Visual Models From Natural Language Supervision
- ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
- FIRST: A Framework for Optimizing Information Quality in Mobile Crowdsensing Systems
Related