vix.ing · top · new · best · stats

Unsupervised Extraction of Video Highlights Via Robust Recurrent Auto-encoders

2015/10/06 by Huan Yang, Baoyuan Wang, Yang, Huan +9 · 26 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Artificial intelligence #Class (philosophy) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Encoder #FOS: Computer and information sciences #Heuristic #Image (mathematics) #Machine learning #Matching (statistics) #Noise (video) #Popularity #Set (abstract data type) #Social media #Unsupervised learning #Video Analysis and Summarization #World Wide Web #cs.CV

paper · pdf · doi:10.48550/arxiv.1510.01442

published in arXiv (Cornell University) (Cornell University) · To Appear in ICCV 2015

arxiv created 2015/10/06 · openalex publication_date 2015/10/06 · arxiv updated 2015/10/07 · openalex created_date 2019/06/27 · openalex updated_date 2026/07/28

Abstract

With the growing popularity of short-form video sharing platforms such as \emInstagram and \emVine, there has been an increasing need for techniques that automatically extract highlights from video. Whereas prior works have approached this problem with heuristic rules or supervised learning, we present an unsupervised learning approach that takes advantage of the abundance of user-edited videos on social media websites such as YouTube. Based on the idea that the most significant sub-events within a video class are commonly present among edited videos while less interesting ones appear less frequently, we identify the significant sub-events via a robust recurrent auto-encoder trained on a collection of user-edited videos queried for each particular class of interest. The auto-encoder is trained using a proposed shrinking exponential loss function that makes it robust to noise in the web-crawled training data, and is configured with bidirectional long short term memory (LSTM)~\citeLSTM:97 cells to better model the temporal structure of highlight segments. Different from supervised techniques, our method can infer highlights using only a set of downloaded edited videos, without also needing their pre-edited counterparts which are rarely available online. Extensive experiments indicate the promise of our proposed solution in this challenging unsupervised settin

Citations

Cited by

Related