HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
2025/12/16 by Ben-Ami, Dan, Serussi, Gabriele, Cohen, Kobi +1
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition
paper · doi:10.48550/arxiv.2512.14870
Abstract
Video Large Language Models (Video-LLMs) are improving rapidly, yet current Video Question Answering (VideoQA) benchmarks often admit single-cue shortcuts, under-testing reasoning that must integrate evidence across time. We introduce HERBench, a benchmark designed to make multi-evidence integration unavoidable: each question requires at least three non-overlapping cues drawn from distinct video segments. HERBench contains 26,806 five-way multiple-choice questions across 12 compositional tasks. To make evidential demand measurable, we introduce the Minimum Required Frame-Set (MRFS), the smallest number of frames a model must fuse to answer correctly, and show that HERBench imposes higher evidential demand than prior benchmarks. Evaluating 13 state-of-the-art Video-LLMs yields only 31-42% accuracy, only modestly above the 20% random-guess baseline. We disentangle this failure into two critical bottlenecks: (1) a retrieval deficit, where frame selectors overlook key evidence, and (2) a fusion deficit, where models fail to integrate information even when all necessary evidence is provided. HERBench thus provides a principled benchmark for studying robust multi-evidence video understanding.
Citations
- Qwen3-VL Technical Report
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Ovis2.5 Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
- Qwen3 Technical Report
- MINERVA: Evaluating Complex Video Reasoning
- T*: Re-thinking Temporal Search for Long-Form Video Understanding
- BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
- Gemma 3 Technical Report
- Adaptive Keyframe Sampling for Long Video Understanding
- Qwen2.5-VL Technical Report
- HD-EPIC: A Highly-Detailed Egocentric Video Dataset
- LLaVA-OneVision: Easy Visual Task Transfer
- The Llama 3 Herd of Models
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- STAR: A Benchmark for Situated Reasoning in Real-World Videos
- TempCompass: Do Video LLMs Really Understand Videos?
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- Can I Trust Your Answer? Visually Grounded Video Question Answering
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Large Scale Real-World Multi-Person Tracking
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Language as Queries for Referring Video Object Segmentation
- End-to-End Referring Video Object Segmentation with Multimodal Transformers
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- TransNet V2: An effective deep network architecture for fast shot transition detection
- CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- TVQA: Localized, Compositional Video Question Answering
- Actor and Action Video Segmentation from a Sentence
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- Simple Online and Realtime Tracking with a Deep Association Metric
- MovieQA: Understanding Stories in Movies through Question-Answering
- Microsoft COCO: Common Objects in Context
Related