EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
2023/08/17 by Karttikeya Mangalam, Mangalam, Karttikeya, Raiymbek Akshulakov +3 · 241 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2308.09126
openalex publication_date 2023/08/17 · openalex created_date 2023/08/22 · openalex updated_date 2026/07/28
Abstract
We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human curated multiple choice question answer pairs, spanning over 250 hours of real video data, covering a very broad range of natural human activity and behavior. For each question, EgoSchema requires the correct answer to be selected between five given options based on a three-minute-long video clip. While some prior works have proposed video datasets with long clip lengths, we posit that merely the length of the video clip does not truly capture the temporal difficulty of the video task that is being considered. To remedy this, we introduce temporal certificate sets, a general notion for capturing the intrinsic temporal understanding length associated with a broad range of video understanding tasks & datasets. Based on this metric, we find EgoSchema to have intrinsic temporal lengths over 5.7x longer than the second closest dataset and 10x to 100x longer than any other video understanding dataset. Further, our evaluation of several current state-of-the-art video and language models shows them to be severely lacking in long-term video understanding capabilities. Even models with several billions of parameters achieve QA accuracy less than 33% (random is 20%) on the EgoSchema multi-choice question answering task, while humans achieve about 76% accuracy. We posit that \name, with its long intrinsic temporal structures and diverse complexity, would serve as a valuable evaluation probe for developing effective long-term video understanding systems in the future. Data and Zero-shot model evaluation code are open-sourced for both public and commercial use under the Ego4D license at http://egoschema.github.io
Cited by
- Video Understanding: From Geometry and Semantics to Unified Models
- WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
- LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
- Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
- KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
- Rethinking Chain-of-Thought Reasoning for Videos
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- What Happens When: Learning Temporal Orders of Events in Videos
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- UniComp: Rethinking Video Compression Through Informational Uniqueness
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Geometrically-Constrained Agent for Spatial Reasoning
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- Video Finetuning Improves Reasoning Between Frames
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
- PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
- MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
- CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- A Video Is Not Worth a Thousand Words
- Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- [De|Re]constructing VLMs' Reasoning in Counting
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
- Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- FrameOracle: Learning What to See and How Much to See in Videos
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- NeMo: Needle in a Montage for Video-Language Understanding
- Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Confidence-guided Refinement Reasoning for Zero-shot Question Answering
- MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
- See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
- MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment
- AdsQA: Towards Advertisement Video Understanding
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- Video-LLMs with Temporal Visual Screening
- CVBench: Evaluating Cross-Video Synergies for Complex Multimodal Understanding and Reasoning
- MovieCORE: COgnitive REasoning in Movies
- An Empirical Study on How Video-LLMs Answer Video Questions
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- Vision Generalist Model: A Survey
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- Episodic Memory Representation for Long-form Video Understanding
- KFFocus: Highlighting Keyframes for Enhanced Video Understanding
- FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Controllable Hybrid Captioner for Improved Long-form Video Understanding
- VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- iSafetyBench: A video-language benchmark for safety in industrial environment
- Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
- ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
- Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
- VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
- Infinite Video Understanding
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
- Scaling RL to Long Videos
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
- ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment
- UVLM: Benchmarking Video Language Model for Underwater World Understanding
- LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
- Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
- LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
- ARGUS: Hallucination and Omission Evaluation in Video-LLMs
- PEVLM: Parallel Encoding for Vision-Language Models
- Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
- LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
- Moment Sampling in Video LLMs for Long-Form Video QA
- FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
- SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
- EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
- ExAct: A Video-Language Benchmark for Expert Action Analysis
- VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
- Movie Facts and Fibs (MF2): A Benchmark for Long Movie Understanding
- EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
- SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
- R3-VQA: "Read the Room" by Video Social Reasoning
- SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
- MiMo-VL Technical Report
- HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
- Seeing the Arrow of Time in Large Multimodal Models
- METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
- EgoVLM: Policy Optimization for Egocentric Video Understanding
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
- VUDG: A Dataset for Video Understanding Domain Generalization
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
- ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- Music's Multimodal Complexity in AVQA: Why We Need More than General Multimodal LLMs
- HoliTom: Holistic Token Merging for Fast Video Large Language Models
- HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
- TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
- Two Causally Related Needles in a Video Haystack
- RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
- ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
- Multimodal Conversation Structure Understanding
- From Evaluation to Defense: Advancing Safety in Video Large Language Models
- Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles
- ViQAgent: Zero-Shot Video Question Answering via Agent with Open-Vocabulary Grounding Validation
- STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
- Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
- Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
- HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
- Understanding Complexity in VideoQA via Visual Program Generation
- Think in Sets for Streaming Video Token Compression
- Visuospatial Cognitive Assistant
- VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
- Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
- EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next
- Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
- SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
- MiniMax Sparse Attention
- Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
- SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
- FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
- MVEB: Massive Video Embedding Benchmark
- EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
- TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
- ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
- VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
- The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
- Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos
- InstructionBench: An Instructional Video Understanding Benchmark
- VideoVista-CulturalLingo: 360^∘ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
- Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
- ViSMaP: Unsupervised Hour-long Video Summarisation by Meta-Prompting
- Advancing Egocentric Video Question Answering with Multimodal Large Language Models
- VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
- Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
- AdaVid: Adaptive Video-Language Pretraining
- HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
- Multimodal Long Video Modeling Based on Temporal Dynamic Context
- VideoAds for Fast-Paced Video Understanding
- PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
- Kimi-VL Technical Report
Related