UCF-101: A dataset of 101 human actions classes from videos in the wild
2012/12/03 by Khurram Soomro, Amir Zamir, Soomro, Khurram +3 · 307 citations
Computer Science · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1212.0402
openalex publication_date 2012/12/03 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/31
Abstract
We introduce UCF101 which is currently the largest dataset of human actions. It consists of 101 action classes, over 13k clips and 27 hours of video data. The database consists of realistic user uploaded videos containing camera motion and cluttered background. Additionally, we provide baseline action recognition results on this new dataset using standard bag of words approach with overall performance of 44.5%. To the best of our knowledge, UCF101 is currently the most challenging dataset of actions due to its large number of classes, large number of clips and also unconstrained nature of such clips.
Cited by
- Autoregressive Flow Matching for Motion Prediction
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- Investigating Deep Learning Models for Ejection Fraction Estimation from Echocardiography Videos
- Test-Time Adaptation via Dual Distillation for Videos Under Severe Distribution Shifts
- SemCovert: Secure and Covert Video Transmission via Deep Semantic-Level Hiding
- OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
- Temporal Interlacing Network
- Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression
- I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models
- Asymmetric Hierarchical Anchoring for Robust Audio-Visual Cross-Modal Generalization
- Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
- Effect of Activation Function and Model Optimizer on the Performance of Human Activity Recognition System Using Various Deep Learning Models
- Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
- Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
- AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
- VICTOR: Dataset Copyright Auditing in Video Recognition Systems
- Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
- Distill Video Datasets into Images
- Adapting MLLMs for Nuanced Video Retrieval
- StegaVAR: Privacy-Preserving Video Action Recognition via Steganographic Domain Analysis
- MetaTPT: Meta Test-time Prompt Tuning for Vision-Language Models
- SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
- Task-Specific Distance Correlation Matching for Few-Shot Action Recognition
- Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
- Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot Learning
- A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques
- Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language Models
- Unleashing Temporal Capacity of Spiking Neural Networks through Spatiotemporal Separation
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
- Heatmap Pooling Network for Action Recognition from RGB Videos
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- Hierarchical Semantic Alignment for Image Clustering
- Partially Shared Concept Bottleneck Models
- Structured Context Learning for Generic Event Boundary Detection
- From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- Seeing without Pixels: Perception from Camera Trajectories
- GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models
- AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
- Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
- VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
- Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
- ReBaPL: Repulsive Bayesian Prompt Learning
- SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
- BoxingVI: A Multi-Modal Benchmark for Boxing Action Recognition and Localization
- VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation
- Decoupling Complexity from Scale in Latent Diffusion Model
- Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning
- Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- Calibrated Multimodal Representation Learning with Missing Modalities
- BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning
- Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
- Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
- Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
- MiVID: Multi-Strategic Self-Supervision for Video Frame Interpolation using Diffusion Model
- Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
- Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
- Wireless Video Semantic Communication with Decoupled Diffusion Multi-frame Compensation
- FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
- Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
- A Retrospect to Multi-prompt Learning across Vision and Language
- Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
- A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models
- When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs
- On the Provable Importance of Gradients for Language-Assisted Image Clustering
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models
- Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models
- Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
- VeFA: Vector-Based Feature Space Adaptation for Robust Model Fine-Tuning
- FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- DGME-T: Directional Grid Motion Encoding for Transformer-Based Historical Camera Movement Classification
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
- Exploring Cross-Modal Flows for Few-Shot Learning
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- BIGFix: Bidirectional Image Generation with Token Fixing
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Mixup Helps Understanding Multimodal Video Better
- Deep Multimodal Feature Encoding for Video Ordering
- microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Perceptron Synthesis Network: Rethinking the Action Scale Variances in Videos
- Cooperative Pseudo Labeling for Unsupervised Federated Classification
- Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
- Temporal Action Detection by Joint Identification-Verification
- D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting Mitigation
- Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
- Few-shot Action Recognition with Implicit Temporal Alignment and Pair Similarity Optimization
- Spatio-Temporal Action Detection with Multi-Object Interaction
- Knowing What, Where and When to Look: Efficient Video Action Modeling\n with Attention
- Temporal Accumulative Features for Sign Language Recognition
- Two-Stream AMTnet for Action Detection
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
- Bridging Text and Video Generation: A Survey
- WebVision Database: Visual Learning and Understanding from Web Data
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Similarity R-C3D for Few-shot Temporal Activity Detection
- From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
- Zero-Shot Decentralized Federated Learning
- SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP
- MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning
- DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
- NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
- Multi-Modal Three-Stream Network for Action Recognition
- Representing Videos as Discriminative Sub-graphs for Action Recognition
- Boosting Video Representation Learning with Multi-Faceted Integration
- Personalizing Pre-trained Models
- Datasets on object manipulation and interaction: a survey
- Deep Multi-Kernel Convolutional LSTM Networks and an Attention-Based\n Mechanism for Videos
- Disentangling Static and Dynamic Information for Reducing Static Bias in Action Recognition
- Adaptive and Iteratively Improving Recurrent Lateral Connections
- Hierarchical Representation Matching for CLIP-based Class-Incremental Learning
- Category Discovery: An Open-World Perspective
- Robust Visual Object Tracking with Two-Stream Residual Convolutional Networks
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- Action Completion: A Temporal Model for Moment Detection
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Synthetic Defocus and Look-Ahead Autofocus for Casual Videography
- DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation
- Shaping Initial State Prevents Modality Competition in Multi-modal Fusion: A Two-stage Scheduling Framework via Fast Partial Information Decomposition
- A Short Note about Kinetics-600
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- CCVS: Context-aware Controllable Video Synthesis
- Body Joint guided 3D Deep Convolutional Descriptors for Action Recognition
- Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition
- Multi-Source Video Domain Adaptation with Temporal Attentive Moment Alignment
- Human Action Recognition and Prediction: A Survey
- Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
- DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition
- Channel Pruning Guided by Classification Loss and Feature Importance
- Out-of-Distribution Detection for Generalized Zero-Shot Action Recognition
- Revisiting hand-crafted feature for action recognition: a set of improved dense trajectories
- Alternating Training-based Label Smoothing Enhances Prompt Generalization
- Context-aware and Scale-insensitive Temporal Repetition Counting
- M-PACT: An Open Source Platform for Repeatable Activity Classification Research
- An Improved Video Analysis using Context based Extension of LSH
- Visual Recognition Using Directional Distribution Distance
- MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
- Is It Certainly a Deepfake? Reliability Analysis in Detection & Generation Ecosystem
- A2M2-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition
- TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
- Discriminative convolutional Fisher vector network for action recognition
- Global Prompt Refinement with Non-Interfering Attention Masking for One-Shot Federated Learning
- MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
- Attentional Pooling for Action Recognition
- \mathttM3VIR: A Large-Scale Multi-Modality Multi-View Synthesized Benchmark Dataset for Image Restoration and Content Creation
- KRAST: Knowledge-Augmented Robotic Action Recognition with Structured Text for Vision-Language Models
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- Self Identity Mapping
- Vision-Language Models for Vision Tasks: A Survey
- Learning Representations from Audio-Visual Spatial Alignment
- Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video\n Action Recognition
- Learning Temporal Embeddings for Complex Video Analysis
- Temporal Attentive Alignment for Large-Scale Video Domain Adaptation
- Regularizing Deep Neural Networks by Noise: Its Interpretation and Optimization
- iMiGUE: An Identity-free Video Dataset for Micro-Gesture Understanding and Emotion Analysis
- Tubelets: Unsupervised action proposals from spatiotemporal super-voxels
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
- AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for\n Human Action Recognition in Videos
- Human Action Forecasting by Learning Task Grammars
- Squeeze-and-Excitation on Spatial and Temporal Deep Feature Space for Action Recognition
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion
- Action Anticipation By Predicting Future Dynamic Images
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Symmetry and Group in Attribute-Object Compositions
- MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
- Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
- Review of Video Predictive Understanding: Early Action Recognition and Future Action Prediction
- Self-Supervised Video Representation Learning with Meta-Contrastive Network
- Label Smoothing++: Enhanced Label Regularization for Training Neural Networks
- REPAIR: Removing Representation Bias by Dataset Resampling
- Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
- AttriPrompt: Dynamic Prompt Composition Learning for CLIP
- DeepStream: Prototyping Deep Joint Source-Channel Coding for Real-Time Multimedia Transmissions
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- Improvements of Motion Estimation and Coding using Neural Networks
- Learning Spatio-Temporal Representation with Local and Global Diffusion
- Actor-Action Semantic Segmentation with Grouping Process Models
- Label Efficient Learning of Transferable Representations across Domains and Tasks
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Real-time Action Recognition with Enhanced Motion Vector CNNs
- On Flow Profile Image for Video Representation
- Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
- Causality-guided Prompt Learning for Vision-language Models via Visual Granulation
- Spatiotemporal Pyramid Network for Video Action Recognition
- Singular Value Few-shot Adaptation of Vision-Language Models
- Towards Efficient General Feature Prediction in Masked Skeleton Modeling
- Deep Spatio-temporal Manifold Network for Action Recognition
- Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison
- FedAPT: Federated Adversarial Prompt Tuning for Vision-Language Models
- Enhancing Fitness Movement Recognition with Attention Mechanism and Pre-Trained Feature Extractors
- Multi-View Intact Space Learning
- Dimensionality Reduction on Grassmannian via Riemannian Optimization: A Generalized Perspective
- Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective
- Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
- Empowering cyberphysical systems of systems with intelligence
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- Self-Supervised Learning via multi-Transformation Classification for Action Recognition
- Sympathy for the Details: Dense Trajectories and Hybrid Classification\n Architectures for Action Recognition
- EventNet: A Large Scale Structured Concept Library for Complex Event Detection in Video
- Spotlighter: Revisiting Prompt Tuning from a Representative Mining View
- Language-Aware Information Maximization for Transductive Few-Shot CLIP
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering
- What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
- Improving performance of recurrent neural network with relu nonlinearity
- Knowledge Fusion Transformers for Video Action Recognition
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep\n Architecture
- Class-Wise Difficulty-Balanced Loss for Solving Class-Imbalance
- Spatiotemporal Contrastive Video Representation Learning
- AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning
- Video Representation Learning with Visual Tempo Consistency
- Hierarchical Contrastive Motion Learning for Video Action Recognition
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian Alignment
- Efficient Multi-Source Knowledge Transfer by Model Merging
- Neural Task Graphs: Generalizing to Unseen Tasks from a Single Video Demonstration
- Human Action Recognition using Factorized Spatio-Temporal Convolutional Networks
- Ontology Based Global and Collective Motion Patterns for Event Classification in Basketball Videos
- Better and Faster: Knowledge Transfer from Multiple Self-supervised Learning Tasks via Graph Distillation for Video Classification
- Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
- Busy-Quiet Video Disentangling for Video Classification
- Actionness Estimation Using Hybrid Fully Convolutional Networks
- Open Set Domain Adaptation for Image and Action Recognition
- A Pursuit of Temporal Accuracy in General Activity Detection
- OmViD: Omni-supervised active learning for video action detection
- Generative Model-Based Feature Attention Module for Video Action Analysis
- Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning
- Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models
- Age of Semantic Information-Aware Wireless Transmission for Remote Monitoring Systems
- Federated Cross-Modal Style-Aware Prompt Generation
- GAN for Vision, KG for Relation: a Two-stage Deep Network for Zero-shot Action Recognition
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
- QuickMerge++: Fast Token Merging with Autoregressive Prior
- AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
- M3OOD: Automatic Selection of Multimodal OOD Detectors
- Recurrent Residual Module for Fast Inference in Videos
- ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- SemPT: Semantic Prompt Tuning for Vision-Language Models
- Few-shot Vision-based Human Activity Recognition with MLLM-based Visual Reinforcement Learning
- A Structured Model For Action Detection
- Human Action Recognition with Multi-Laplacian Graph Convolutional Networks
- Multi-Scale Video Frame-Synthesis Network with Transitive Consistency Loss
- Relational Action Forecasting
- Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- AME: Aligned Manifold Entropy for Robust Vision-Language Distillation
- Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
- Stacked dense optical flows and dropout layers to predict sperm motility and morphology
- Objects2action: Classifying and localizing actions without any video\n example
- Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
- Self-Supervised Representation Learning for Visual Anomaly Detection
- BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models
- MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
- Multimodal learning with next-token prediction for large multimodal models
- Dynamic Concept Composition for Zero-Example Event Detection
- Feature-Supervised Action Modality Transfer
- The Trauma THOMPSON Dataset for Real-World Emergency AI
- Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise
- Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
- Do Less and Achieve More: Training CNNs for Action Recognition Utilizing Action Images from the Web
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
- MoExDA: Domain Adaptation for Edge-based Action Recognition
- Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
- Attention is all you need for Videos: Self-attention based Video\n Summarization using Universal Transformers
- Raw Data Matters: Enhancing Prompt Tuning by Internal Augmentation on Vision-Language Models
- EvoVLMA: Evolutionary Vision-Language Model Adaptation
- ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models
- FASTER Recurrent Networks for Efficient Video Classification
- Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models
- Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
- Initialization Strategies of Spatio-Temporal Convolutional Neural\n Networks
- RainbowPrompt: Diversity-Enhanced Prompt-Evolving for Continual Learning
- GVD: Guiding Video Diffusion Model for Scalable Video Distillation
- MOVE: Motion-Guided Few-Shot Video Object Segmentation
- Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
Related