NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
2021/05/18 by Junbin Xiao, Xiao, Junbin, Xindi Shang +5 · 126 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Domain Adaptation and Few-Shot Learning
paper · pdf · doi:10.48550/arxiv.2105.08276
Abstract
We introduce NExT-QA, a rigorously designed video question answering (VideoQA) benchmark to advance video understanding from describing to explaining the temporal actions. Based on the dataset, we set up multi-choice and open-ended QA tasks targeting causal action reasoning, temporal action reasoning, and common scene comprehension. Through extensive analysis of baselines and established VideoQA techniques, we find that top-performing methods excel at shallow scene descriptions but are weak in causal and temporal action reasoning. Furthermore, the models that are effective on multi-choice QA, when adapted to open-ended QA, still struggle in generalizing the answers. This raises doubt on the ability of these models to reason and highlights possibilities for improvement. With detailed results for different question types and heuristic observations for future works, we hope NExT-QA will guide the next generation of VQA research to go beyond superficial scene description towards a deeper understanding of videos. (The dataset and related resources are available at https://github.com/doc-doc/NExT-QA.git)
Citations
Cited by
- LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- Selective LoRA for Visual Tokens and Attention Heads
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- Adapting MLLMs for Nuanced Video Retrieval
- HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
- VisualActBench: Can VLMs See and Act like a Human?
- Towards Lossless Ultimate Vision Token Compression for VLMs
- A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
- PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- COACH: Collaborative Agents for Contextual Highlighting -- A Multi-Agent Framework for Sports Video Analysis
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models
- VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
- Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
- AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
- OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
- ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- NVIDIA Nemotron Nano V2 VL
- When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning
- Spot The Ball: A Benchmark for Visual Social Inference
- LongCat-Flash-Omni Technical Report
- Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
- VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
- Large Emotional World Model
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- A Video Is Not Worth a Thousand Words
- Positional Preservation Embedding for Multimodal Large Language Models
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- Not in Sync: Unveiling Temporal Bias in Audio Chat Models
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- VideoNorms: Benchmarking Cultural Awareness of Video Language Models
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- Addressing the ID-Matching Challenge in Long Video Captioning
- LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- FrameOracle: Learning What to See and How Much to See in Videos
- Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
- POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
- NeMo: Needle in a Montage for Video-Language Understanding
- Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
- Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
- FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
- SPIKE-RL: Video-LLMs meet Bayesian Surprise
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning
- Confidence-guided Refinement Reasoning for Zero-shot Question Answering
- See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
- Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
- Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
- AToken: A Unified Tokenizer for Vision
- SAIL-VL2 Technical Report
- Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
- Video Understanding by Design: How Datasets Shape Architectures and Insights
- AdsQA: Towards Advertisement Video Understanding
- Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
- Do Video Language Models Really Know Where to Look? Diagnosing Attention Failures in Video Language Models
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
- ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- Video-LLMs with Temporal Visual Screening
- CVBench: Evaluating Cross-Video Synergies for Complex Multimodal Understanding and Reasoning
- SpotEdit: Evaluating Visually-Guided Image Editing Methods
- Mitigating Easy Option Bias in Multiple-Choice Question Answering
- Causality Matters: How Temporal Information Emerges in Video Language Models
- CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
- ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Object-centric Video Question Answering with Visual Grounding and Referring
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Related