The Kinetics Human Action Video Dataset
2017/05/19 by Will Kay, João Carreira, Kay, Will +21 · 240 citations
Computer Science · #Human Pose and Action Recognition #Anomaly Detection Techniques and Applications #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.1705.06950
Abstract
We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 10s and is taken from a different YouTube video. The actions are human focussed and cover a broad range of classes including human-object interactions such as playing instruments, as well as human-human interactions such as shaking hands. We describe the statistics of the dataset, how it was collected, and give some baseline performance figures for neural network architectures trained and tested for human action classification on this dataset. We also carry out a preliminary analysis of whether imbalance in the dataset leads to bias in the classifiers.
Cited by
- Autoregressive Flow Matching for Motion Prediction
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- LangPrecip: Language-Aware Multimodal Precipitation Nowcasting
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Open Your Model’s Eyes: Video and Context-Aware Multimodal Backchannel Prediction
- FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video Editing
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
- Non-Contrast CT Esophageal Varices Grading through Clinical Prior-Enhanced Multi-Organ Analysis
- Hierarchical Bayesian Framework for Multisource Domain Adaptation
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- VICTOR: Dataset Copyright Auditing in Video Recognition Systems
- Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- Recurrent Video Masked Autoencoders
- Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces
- MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
- VisualActBench: Can VLMs See and Act like a Human?
- Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
- Opinion: Learning Intuitive Physics May Require More than Visual Data
- Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
- Detection of Intoxicated Individuals from Facial Video Sequences via a Recurrent Fusion Model
- Unique Lives, Shared World: Learning from Single-Life Videos
- Heatmap Pooling Network for Action Recognition from RGB Videos
- ProtoEFNet: Dynamic Prototype Learning for Inherently Interpretable Ejection Fraction Estimation in Echocardiography
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- Structured Context Learning for Generic Event Boundary Detection
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- DisMo: Disentangled Motion Representations for Open-World Motion Transfer
- SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
- GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models
- Smooth regularization for efficient video recognition
- LungEvaty: A Scalable, Open-Source Transformer-based Deep Learning Model for Lung Cancer Risk Prediction in LDCT Screening
- VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
- Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
- Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- BoxingVI: A Multi-Modal Benchmark for Boxing Action Recognition and Localization
- MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
- RoCoISLR: A Romanian Corpus for Isolated Sign Language Recognition
- Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System
- RodEpil: A Video Dataset of Laboratory Rodents for Seizure Detection and Benchmark Evaluation
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- RadHARSimulator V2: Video to Doppler Generator
- FAST-CAD: A Fairness-Aware Framework for Non-Contact Stroke Diagnosis
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
- Balancing Multi-modal Sensor Learning via Multi-objective Optimization
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
- RF-Behavior: A Multimodal Radio-Frequency Dataset for Human Behavior and Emotion Analysis
- Temporal Zoom Networks: Distance Regression and Continuous Depth for Efficient Action Localization
- Web-Scale Collection of Video Data for 4D Animal Reconstruction
- Dynamic Reflections: Probing Video Representations with Text Alignment
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual Decoding
- FastBoost: Progressive Attention with Dynamic Scaling for Efficient Deep Learning
- Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
- Pyramidal Convolution: Rethinking Convolutional Neural Networks for Visual Recognition
- Understanding Human Hands in Contact at Internet Scale
- Gated Channel Transformation for Visual Recognition
- Anomaly Locality in Video Surveillance
- Learning from Temporal Gradient for Semi-supervised Action Recognition
- Survey: Transformer based Video-Language Pre-training
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Action-Sufficient State Representation Learning for Control with Structural Constraints
- Egok360: A 360 Egocentric Kinetic Human Activity Video Dataset
- Blindly Assess Quality of In-the-Wild Videos via Quality-Aware Pre-Training and Motion Perception
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Play Fair: Frame Attributions in Video Models
- Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution
- Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
- AEI: Actors-Environment Interaction with Adaptive Attention for Temporal Action Proposals Generation
- TSI: Temporal Saliency Integration for Video Action Recognition
- Video Modeling with Correlation Networks
- Video Understanding as Machine Translation
- ActBERT: Learning Global-Local Video-Text Representations
- A Real-time Action Representation with Temporal Encoding and Deep Compression
- Motion Feature Network: Fixed Motion Filter for Action Recognition
- Study of Spatio-Temporal Modeling in Video Quality Assessment
- An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos
- Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
- Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey
- Efficient Spatialtemporal Context Modeling for Action Recognition
- Self-Supervised Video Representation Learning with Space-Time Cubic Puzzles
- Temporal Bilinear Networks for Video Action Recognition
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency Detection
- Temporal Gaussian Mixture Layer for Videos
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QA
- PANDA: A Gigapixel-level Human-centric Video Dataset
- Video Action Recognition Via Neural Architecture Searching
- A Self Validation Network for Object-Level Human Attention Estimation
- Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification
- Diagnosing Error in Temporal Action Detectors
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- Cycle-Contrast for Self-Supervised Video Representation Learning
- The AVA-Kinetics Localized Human Actions Video Dataset
- TAda! Temporally-Adaptive Convolutions for Video Understanding
- Detecting Attended Visual Targets in Video
- Deep Analysis of CNN-based Spatio-temporal Representations for Action Recognition
- Weakly Supervised Action Selection Learning in Video
- Spatio-Temporal Graph for Video Captioning with Knowledge Distillation
- Revisiting Temporal Modeling for Video-based Person ReID
- GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
- BASAR:Black-box Attack on Skeletal Action Recognition
- Temporal Action Localization with Variance-Aware Networks
- Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
- FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- Swin Transformer V2: Scaling Up Capacity and Resolution
- Quantifying Multimodal Imbalance: A GMM-Guided Adaptive Loss for Audio-Visual Learning
- A Comprehensive Survey on World Models for Embodied AI
- Thinking in Frequency: Face Forgery Detection by Mining Frequency-aware Clues
- Video Swin Transformer
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- Spatial-Temporal Alignment Network for Action Recognition and Detection
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
- Mixup Helps Understanding Multimodal Video Better
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer
- Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
- Class-Aware Adversarial Lung Nodule Synthesis in CT Images
- Multi-Object Tracking with Hallucinated and Unlabeled Videos
- Low Pass Filter for Anti-aliasing in Temporal Action Localization
- Multi-object tracking with self-supervised associating network
- Few-shot Action Recognition with Implicit Temporal Alignment and Pair Similarity Optimization
- G-TAD: Sub-Graph Localization for Temporal Action Detection
- Temporal Accumulative Features for Sign Language Recognition
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Online Generic Event Boundary Detection
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- SimMIM: A Simple Framework for Masked Image Modeling
- NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
- TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
- Deep Multi-Kernel Convolutional LSTM Networks and an Attention-Based\n Mechanism for Videos
- Disentangling Static and Dynamic Information for Reducing Static Bias in Action Recognition
- Adaptive and Iteratively Improving Recurrent Lateral Connections
- Category Discovery: An Open-World Perspective
- VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- Every Subtlety Counts: Fine-grained Person Independence Micro-Action Recognition via Distributionally Robust Optimization
- Application of Transfer Learning to Sign Language Recognition using an\n Inflated 3D Deep Convolutional Neural Network
- Transformer-based audio-visual multimodal fusion for fine-grained recognition of individual sow nursing behaviour
- A Short Note about Kinetics-600
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Look, Listen and Learn
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition
- Multi-Source Video Domain Adaptation with Temporal Attentive Moment Alignment
- Human Action Recognition and Prediction: A Survey
- Fast Video Shot Transition Localization with Deep Structured Models
- DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
- ExtremeWeather: A large-scale climate dataset for semi-supervised\n detection, localization, and understanding of extreme weather events
- RareAct: A video dataset of unusual interactions
- MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
- Attentional Pooling for Action Recognition
- KRAST: Knowledge-Augmented Robotic Action Recognition with Structured Text for Vision-Language Models
- Submission to ActivityNet Challenge 2019: Task B Spatio-temporal Action Localization
- ResidualViT for Efficient Temporally Dense Video Encoding
- Revisiting ResNets: Improved Training and Scaling Strategies
- More performant and scalable: Rethinking contrastive vision-language pre-training of radiology in the LLM era
- Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
- Learning Representations from Audio-Visual Spatial Alignment
- Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video\n Action Recognition
- Temporal Attentive Alignment for Large-Scale Video Domain Adaptation
- MoViNets: Mobile Video Networks for Efficient Video Recognition
- A client–server based recognition system: Non-contact single/multiple emotional and behavioral state assessment methods
- Learning Graph Convolutional Network for Skeleton-based Human Action Recognition by Neural Searching
- Weakly-supervised Fingerspelling Recognition in British Sign Language\n Videos
- Recognizing American Sign Language Manual Signs from RGB-D Videos
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- Video Understanding by Design: How Datasets Shape Architectures and Insights
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Actor-Centric Relation Network
- Self-Supervised Video Representation Learning with Meta-Contrastive Network
- REPAIR: Removing Representation Bias by Dataset Resampling
- Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection
- Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- Fine-grained Video Categorization with Redundancy Reduction Attention
- DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-based Human Action Segmentation
- A Multigrid Method for Efficiently Training Video Models
- Space-time Mixing Attention for Video Transformer
- Efficient Video Transformers with Spatial-Temporal Token Selection
- On Flow Profile Image for Video Representation
- VideoSSL: Semi-Supervised Learning for Video Classification
- Disentangled Non-Local Neural Networks
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion
- Empowering cyberphysical systems of systems with intelligence
- Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering
- What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
- Knowledge Fusion Transformers for Video Action Recognition
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- Spatiotemporal Contrastive Video Representation Learning
- AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning
- Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
- Open Set Domain Adaptation for Image and Action Recognition
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information Bottleneck
- Energy-based Periodicity Mining with Deep Features for Action Repetition Counting in Unconstrained Videos
- Generic Event Boundary Detection via Denoising Diffusion
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
- Recurrent Residual Module for Fast Inference in Videos
- Versatile Video Tokenization with Generative 2D Gaussian Splatting
- ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
- DIVA-VQA: Detecting Inter-frame Variations in UGC Video Quality
- Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
- Improving Automated COVID-19 Grading with Convolutional Neural Networks in Computed Tomography Scans: An Ablation Study
- DeeperForensics Challenge 2020 on Real-World Face Forgery Detection: Methods and Results
- Relational Action Forecasting
- Learning Explicit and Implicit Latent Common Spaces for Audio-Visual Cross-Modal Retrieval
- VGGSounder: Audio-Visual Evaluations for Foundation Models
- Towards Training Stronger Video Vision Transformers for EPIC-KITCHENS-100 Action Recognition
- Self-supervised Learning for Video Correspondence Flow
- The Trauma THOMPSON Dataset for Real-World Emergency AI
- Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
- CRAM: Large-scale Video Continual Learning with Bootstrapped Compression
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
- MoExDA: Domain Adaptation for Edge-based Action Recognition
- Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
- Attention is all you need for Videos: Self-attention based Video\n Summarization using Universal Transformers
- FASTER Recurrent Networks for Efficient Video Classification
- SGCap: Decoding Semantic Group for Zero-shot Video Captioning
- Initialization Using Perlin Noise for Training Networks with a Limited Amount of Data
- MOVE: Motion-Guided Few-Shot Video Object Segmentation
- StepAL: Step-aware Active Learning for Cataract Surgical Videos
- Dynamic GCN: Context-enriched Topology Learning for Skeleton-based Action Recognition
Related