vix.ing · top · new · best · stats · spec

HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips

2019/06/07 by Antoine Miech, Dimitri Zhukov, Miech, Antoine +9 · 153 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization

paper · pdf · doi:10.48550/arxiv.1906.03327

openalex publication_date 2019/06/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30

Abstract

Learning text-video embeddings usually requires a dataset of video clips with\nmanually provided captions. However, such datasets are expensive and time\nconsuming to create and therefore difficult to obtain on a large scale. In this\nwork, we propose instead to learn such embeddings from video data with readily\navailable natural language annotations in the form of automatically transcribed\nnarrations. The contributions of this work are three-fold. First, we introduce\nHowTo100M: a large-scale dataset of 136 million video clips sourced from 1.22M\nnarrated instructional web videos depicting humans performing and describing\nover 23k different visual tasks. Our data collection procedure is fast,\nscalable and does not require any additional manual annotation. Second, we\ndemonstrate that a text-video embedding trained on this data leads to\nstate-of-the-art results for text-to-video retrieval and action localization on\ninstructional video datasets such as YouCook2 or CrossTask. Finally, we show\nthat this embedding transfers well to other domains: fine-tuning on generic\nYoutube videos (MSR-VTT dataset) and movies (LSMDC dataset) outperforms models\ntrained on these datasets alone. Our dataset, code and models will be publicly\navailable at: www.di.ens.fr/willow/research/howto100m/.\n

Citations

Cited by

Related