TALL: Temporal Activity Localization via Language Query
2017/05/05 by Jiyang Gao, Chen Sun, Gao, Jiyang +5 · 104 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.1705.02101
Abstract
This paper focuses on temporal localization of actions in untrimmed videos. Existing methods typically train classifiers for a pre-defined list of actions and apply them in a sliding window fashion. However, activities in the wild consist of a wide combination of actors, actions and objects; it is difficult to design a proper activity list that meets users' needs. We propose to localize activities by natural language queries. Temporal Activity Localization via Language (TALL) is challenging as it requires: (1) suitable design of text and video representations to allow cross-modal matching of actions and language queries; (2) ability to locate actions accurately given features from sliding windows of limited granularity. We propose a novel Cross-modal Temporal Regression Localizer (CTRL) to jointly model text query and video clips, output alignment scores and action boundary regression results for candidate clips. For evaluation, we adopt TaCoS dataset, and build a new dataset for this task on top of Charades by adding sentence temporal annotations, called Charades-STA. We also build complex sentence queries in Charades-STA for test. Experimental results show that CTRL outperforms previous methods significantly on both datasets.
Cited by
- TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
- The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Streaming Video Instruction Tuning
- Object-Centric Framework for Video Moment Retrieval
- Xiaomi MiMo-VL-Miloco Technical Report
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- Moment and Highlight Detection via MLLM Frame Segmentation
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Point to Span: Zero-Shot Moment Retrieval for Navigating Unseen Hour-Long Videos
- ChronusOmni: Improving Time Awareness of Omni Large Language Models
- START: Spatial and Textual Learning for Chart Understanding
- Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- OneThinker: All-in-one Reasoning Model for Image and Video
- IVCR-200K: A Large-Scale Multi-turn Dialogue Benchmark for Interactive Video Corpus Retrieval
- Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
- See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- Qwen3-VL Technical Report
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
- Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training
- SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
- Video Finetuning Improves Reasoning Between Frames
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict
- Mitigating Semantic Collapse in Partially Relevant Video Retrieval
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- Motion-Appearance Co-Memory Networks for Video Question Answering
- RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
- VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
- HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
- Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding
- Augmenting Moment Retrieval: Zero-Dependency Two-Stage Learning
- When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- An empirical study of the effect of video encoders on Temporal Video Grounding
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval
- Not in Sync: Unveiling Temporal Bias in Audio Chat Models
- SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
- Improving Temporal Understanding Logic Consistency in Video-Language Models via Attention Enhancement
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- NeMo: Needle in a Montage for Video-Language Understanding
- Sim-DETR: Unlock DETR for Temporal Sentence Grounding
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
- AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
- ResidualViT for Efficient Temporally Dense Video Encoding
- Rudder: A Cross Lingual Video and Text Retrieval Dataset
- Exploiting Temporal Relationships in Video Moment Localization with Natural Language
- Video Understanding by Design: How Datasets Shape Architectures and Insights
- Harnessing Object Grounding for Time-Sensitive Video Understanding
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language
- On Pursuit of Designing Multi-modal Transformer for Video Grounding
- Video-LLMs with Temporal Visual Screening
- ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval
- Aligning Moments in Time using Video Queries
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- OVG-HQ: Online Video Grounding with Hybrid-modal Queries
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
- TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
- Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks for Enhanced Action Understanding
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- Length Matters: Length-Aware Transformer for Temporal Sentence Grounding
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
- Video Moment Retrieval via Natural Language Queries
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
- LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
- HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
- SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities
- VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning
- Sparse-Dense Side-Tuner for efficient Video Temporal Grounding
- Online Detection of Action Start in Untrimmed, Streaming Videos
- Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- Boosting Temporal Sentence Grounding via Causal Inference
Related