MLVU: Benchmarking Multi-task Long Video Understanding
2024/06/06 by Junjie Zhou, Yan Shu, Zhou, Junjie +19 · 128 citations
Computer Science · #Human Pose and Action Recognition #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.2406.04264
Abstract
The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multi-task Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: 1) The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. 2) The inclusion of various video genres, e.g., movies, surveillance footage, egocentric videos, cartoons, game videos, etc., which reflects the models' LVU performances in different scenarios. 3) The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs' key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today's technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of long video understanding by providing a comprehensive and in-depth analysis of MLLMs.
Cited by
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
- TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
- MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
- LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
- SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
- Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
- HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- Rethinking Chain-of-Thought Reasoning for Videos
- Towards Lossless Ultimate Vision Token Compression for VLMs
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- What Happens When: Learning Temporal Orders of Events in Videos
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- Qwen3.5-Omni Technical Report
- VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
- UniComp: Rethinking Video Compression Through Informational Uniqueness
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- Qwen3-VL Technical Report
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
- Test-Time Temporal Sampling for Efficient MLLM Video Understanding
- VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
- V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
- SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
- Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
- FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- VideoLucy: Deep Memory Backtracking for Long Video Understanding
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- VideoNSA: Native Sparse Attention Scales Video Understanding
- From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
- VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- SD-MVSum: Script-Driven Multimodal Video Summarization Method and Datasets
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- FrameOracle: Learning What to See and How Much to See in Videos
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
- NeMo: Needle in a Montage for Video-Language Understanding
- FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
- Video Panels for Long Video Understanding
- Evaluating point-light biological motion in multimodal large language models
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
- Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
- VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- Qwen3-Omni Technical Report
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- AToken: A Unified Tokenizer for Vision
- CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
- Do Video Language Models Really Know Where to Look? Diagnosing Attention Failures in Video Language Models
- SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- Video-LLMs with Temporal Visual Screening
- Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
- MovieCORE: COgnitive REasoning in Movies
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Ovis2.5 Technical Report
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- KFFocus: Highlighting Keyframes for Enhanced Video Understanding
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- A Metric for MLLM Alignment in Large-scale Recommendation
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
- Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
- ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
Related