HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips
2019/06/07 by Antoine Miech, Dimitri Zhukov, Miech, Antoine +9 · 153 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.1906.03327
openalex publication_date 2019/06/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Learning text-video embeddings usually requires a dataset of video clips with\nmanually provided captions. However, such datasets are expensive and time\nconsuming to create and therefore difficult to obtain on a large scale. In this\nwork, we propose instead to learn such embeddings from video data with readily\navailable natural language annotations in the form of automatically transcribed\nnarrations. The contributions of this work are three-fold. First, we introduce\nHowTo100M: a large-scale dataset of 136 million video clips sourced from 1.22M\nnarrated instructional web videos depicting humans performing and describing\nover 23k different visual tasks. Our data collection procedure is fast,\nscalable and does not require any additional manual annotation. Second, we\ndemonstrate that a text-video embedding trained on this data leads to\nstate-of-the-art results for text-to-video retrieval and action localization on\ninstructional video datasets such as YouCook2 or CrossTask. Finally, we show\nthat this embedding transfers well to other domains: fine-tuning on generic\nYoutube videos (MSR-VTT dataset) and movies (LSMDC dataset) outperforms models\ntrained on these datasets alone. Our dataset, code and models will be publicly\navailable at: www.di.ens.fr/willow/research/howto100m/.\n
Citations
Cited by
- Autoregressive Flow Matching for Motion Prediction
- The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- How Much 3D Do Video Foundation Models Encode?
- LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
- OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
- TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
- Recurrent Video Masked Autoencoders
- Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
- Opinion: Learning Intuitive Physics May Require More than Visual Data
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- IVCR-200K: A Large-Scale Multi-turn Dialogue Benchmark for Interactive Video Corpus Retrieval
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- Learning Skill-Attributes for Transferable Assessment in Video
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
- CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- Web-Scale Collection of Video Data for 4D Animal Reconstruction
- SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration
- A Step Toward World Models: A Survey on Robotic Manipulation
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- VC4VG: Optimizing Video Captions for Text-to-Video Generation
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- Training-free Online Video Step Grounding
- A Comprehensive Survey on World Models for Embodied AI
- Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Vision Language Models: A Survey of 26K Papers
- ACMID: Automatic Curation of Musical Instrument Dataset for 7-Stem Music Source Separation
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- Bridging Text and Video Generation: A Survey
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Learning Temporal Dynamics from Cycles in Narrated Video
- SSAN: Separable Self-Attention Network for Video Representation Learning
- MASH: A Multiplatform and Multimodal Annotated Dataset for Societal Impact of Hurricane
- Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
- EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- A CLIP-Enhanced Method for Video-Language Understanding
- RareAct: A video dataset of unusual interactions
- Prompt-Driven Agentic Video Editing System: Autonomous Comprehension of Long-Form, Story-Driven Media
- EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
- RLBind: Adversarial-Invariant Cross-Modal Alignment for Unified Robust Embeddings
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
- QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
- Video Understanding by Design: How Datasets Shape Architectures and Insights
- Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
- Multi-modal Transformer for Video Retrieval
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Planning with Reasoning using Vision Language World Model
- SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
- GEM: A General Evaluation Benchmark for Multimodal Tasks
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with\n Transformers
- What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
- The End-of-End-to-End: A Video Understanding Pentathlon Challenge (2020)
- Spatiotemporal Contrastive Video Representation Learning
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
- Cascaded Multilingual Audio-Visual Learning from Videos
- TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Expert Training: Task Hardness Aware Meta-Learning for Few-Shot Classification
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
- Supervision Levels Scale (SLS)
- Unsupervised Discovery of Actions in Instructional Videos
- Back to the Features: DINO as a Foundation for Video World Models
- VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
- Enhancing Scene Transition Awareness in Video Generation via Post-Training
- Learning Video Representations using Contrastive Bidirectional Transformer
- SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval
- HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
- Learning Object Manipulation Skills via Approximate State Estimation from Real Videos
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
- Data Transformation Strategies to Remove Heterogeneity
- DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
- T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval
- Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset
- UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
- Simplifying Traffic Anomaly Detection with Video Foundation Models
- Multi-Granularity Network with Modal Attention for Dense Affective Understanding
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- M3-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
- Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
- Induce, Edit, Retrieve: Language Grounded Multimodal Schema for Instructional Video Retrieval
- UVLM: Benchmarking Video Language Model for Underwater World Understanding
- CI-VID: A Coherent Interleaved Text-Video Dataset
- Modeling short visual events through the BOLD moments video fMRI dataset and metadata
- Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
- Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
- What is More Likely to Happen Next? Video-and-Language Future Event Prediction
- CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
- Multimodal neural networks better explain multivoxel patterns in the\n hippocampus
- End-to-End Learning of Visual Representations from Uncurated Instructional Videos
- Dual Perspectives on Non-Contrastive Self-Supervised Learning
- Can Vision Language Models Understand Mimed Actions?
- EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
- TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
- UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
- Omni-sourced Webly-supervised Learning for Video Recognition
- Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
- ExAct: A Video-Language Benchmark for Expert Action Analysis
- Unleashing Hour-Scale Video Training for Long Video-Language Understanding
- VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations
- Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows
- Video-Text Pre-training with Learned Regions
- VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
- Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
- Vid2Coach: Transforming How-To Videos into Task Assistants
- Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
- Unsupervised Transcript-assisted Video Summarization and Highlight Detection
- PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
- What Do Latent Action Models Actually Learn?
- Weak Supervision and Referring Attention for Temporal-Textual Association Learning
- Video-aided Unsupervised Grammar Induction
- The Role of Video Generation in Enhancing Data-Limited Action Understanding
- Routing with Self-Attention for Multimodal Capsule Networks
- ICYM2I: The illusion of multimodal informativeness under missingness
- Multimodal Pretraining for Dense Video Captioning
- ZR-2021VG: Zero-Resource Speech Challenge, Visually-Grounded Language Modelling track, 2021 edition
- LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Text embedding models can be great data engineers
- Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics
- Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval
- ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- ViSeRet: A simple yet effective approach to moment retrieval via fine-grained video segmentation
Related