Video Understanding by Design: How Datasets Shape Architectures and Insights
2025/09/11 by Wang, Lei, Koniusz, Piotr, Gao, Yongsheng
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2509.09151
Abstract
Video understanding has advanced rapidly, fueled by increasingly complex datasets and powerful architectures. Yet existing surveys largely classify models by task or family, overlooking the structural pressures through which datasets guide architectural evolution. This survey is the first to adopt a dataset-driven perspective, showing how motion complexity, temporal span, hierarchical composition, and multimodal richness impose inductive biases that models should encode. We reinterpret milestones, from two-stream and 3D CNNs to sequential, transformer, and multimodal foundation models, as concrete responses to these dataset-driven pressures. Building on this synthesis, we offer practical guidance for aligning model design with dataset invariances while balancing scalability and task demands. By unifying datasets, inductive biases, and architectures into a coherent framework, this survey provides both a comprehensive retrospective and a prescriptive roadmap for advancing general-purpose video understanding.
Citations
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
- Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
- Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
- MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
- Motion meets Attention: Video Motion Prompts
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
- Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
- Foundation Models for Video Understanding: A Survey
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
- Learning Correlation Structures for Vision Transformers
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- VideoMamba: State Space Model for Efficient Video Understanding
- Sora as a World Model? A Complete Survey on Text-to-Video Generation
- Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
- VideoPrism: A Foundational Visual Encoder for Video Understanding
- Mamba-ND: Selective State Space Modeling for Multi-Dimensional Data
- Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal-Viewpoint Alignment
- Taylor Videos for Action Recognition
- GroundingGPT:Language Enhanced Multi-modal Grounding Model
- Video Understanding with Large Language Models: A Survey
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
- A Survey on Video Diffusion Models
- An Outlook into the Future of Egocentric Vision
- Valley: Video Assistant with Large Language model Enhanced abilitY
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
- VideoChat: Chat-Centric Video Understanding
- End-to-End Spatio-Temporal Action Localisation with Video Transformers
- Search-Map-Search: A Frame Selection Paradigm for Action Recognition
- Diffusion Action Segmentation
- VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
- Unmasked Teacher: Towards Training-Efficient Video Foundation Models
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- Selective Structured State-Spaces for Long-Form Video Understanding
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- Epic-Sounds: A Large-scale Dataset of Actions That Sound
- VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners
- InternVideo: General Video Foundation Models via Generative and Discriminative Learning
- Vision Transformers for Action Recognition: A Survey
- Zero-Shot Video Question Answering via Frozen Bidirectional Language Models
- Multimodal Learning with Transformers: A Survey
- Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
- Masked Autoencoders As Spatiotemporal Learners
- Long Movie Clip Classification with State-Space Video Models
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- All in One: Exploring Unified Video-Language Pre-training
- Omnivore: A Single Model for Many Visual Modalities
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition
- Masked Feature Prediction for Self-Supervised Visual Pre-Training
- MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- ActionCLIP: A New Paradigm for Video Action Recognition
- Video Swin Transformer
- Video Swin Transformer
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- Anticipative Video Transformer
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- SportsCap: Monocular 3D Human Motion Capture and Fine-grained Understanding in Challenging Sports Videos
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Multiscale Vision Transformers
- CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
- UAV-Human: A Large Benchmark for Human Behavior Understanding with Unmanned Aerial Vehicles
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning
- ViViT: A Video Vision Transformer
- MoViNets: Mobile Video Networks for Efficient Video Recognition
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- Is Space-Time Attention All You Need for Video Understanding?
- Video Transformer Network
- Transformers in Vision: A Survey
- Video Generative Adversarial Networks: A Review
- Spatiotemporal Contrastive Video Representation Learning
- Multi-modal Transformer for Video Retrieval
- Self-supervised Learning: Generative or Contrastive
- The AVA-Kinetics Localized Human Actions Video Dataset
- VGGSound: A Large-scale Audio-Visual Dataset
- FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding
- X3D: Expanding Architectures for Efficient Video Recognition
- Temporal Pyramid Network for Action Recognition
- A Survey on 3D Skeleton-Based Action Recognition Using Learning Method
- Temporal Interlacing Network
- Gate-Shift Networks for Video Action Recognition
- Tiny Video Networks
- CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action\n Recognition
- Video Modeling with Correlation Networks
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips
- Attentive Spatio-Temporal Representation Learning for Diving\n Classification
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- VideoBERT: A Joint Model for Video and Language Representation Learning
- Cross-task weakly supervised learning from instructional videos
- MS-TCN: Multi-Stage Temporal Convolutional Network for Action Segmentation
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey
- SlowFast Networks for Video Recognition
- Video Action Transformer Network
- Timeception for Complex Action Recognition
- TSM: Temporal Shift Module for Efficient Video Understanding
- TVQA: Localized, Compositional Video Question Answering
- Videos as Space-Time Region Graphs
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction
- Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition
- Moments in Time Dataset: one million videos for event understanding
- Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- Temporal Relational Reasoning in Videos
- Localizing Moments in Video with Natural Language
- The "something something" video database for learning and evaluating visual common sense
- AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions
- Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
- The Kinetics Human Action Video Dataset
- TALL: Temporal Activity Localization via Language Query
- Dense-Captioning Events in Videos
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding
- Need for Speed: A Benchmark for Higher Frame Rate Object Tracking
- Towards Automatic Learning of Procedures from Web Instructional Videos
- Deep Reinforcement Learning: An Overview
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
- Hierarchical Deep Temporal Models for Group Activity Recognition
- Going Deeper into Action Recognition: A Survey
- Movie Description
- NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis
- Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
- Human Action Recognition using Factorized Spatio-Temporal Convolutional Networks
- Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
- Beyond Short Snippets: Deep Networks for Video Classification
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Cross-view Action Modeling, Learning and Recognition
- Learning Human Activities and Object Affordances from RGB-D Videos
Related