2020/07/28 by Vishal Kaushal, Suraj Kothawade, Kaushal, Vishal +5
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.2007.14560
openalex publication_date 2020/07/28 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Automatic video summarization is still an unsolved problem due to several\nchallenges. We take steps towards making automatic video summarization more\nrealistic by addressing them. Firstly, the currently available datasets either\nhave very short videos or have few long videos of only a particular type. We\nintroduce a new benchmarking dataset VISIOCITY which comprises of longer videos\nacross six different categories with dense concept annotations capable of\nsupporting different flavors of video summarization and can be used for other\nvision problems. Secondly, for long videos, human reference summaries are\ndifficult to obtain. We present a novel recipe based on pareto optimality to\nautomatically generate multiple reference summaries from indirect ground truth\npresent in VISIOCITY. We show that these summaries are at par with human\nsummaries. Thirdly, we demonstrate that in the presence of multiple ground\ntruth summaries (due to the highly subjective nature of the task), learning\nfrom a single combined ground truth summary using a single loss function is not\na good idea. We propose a simple recipe VISIOCITY-SUM to enhance an existing\nmodel using a combination of losses and demonstrate that it beats the current\nstate of the art techniques when tested on VISIOCITY. We also show that a\nsingle measure to evaluate a summary, as is the current typical practice, falls\nshort. We propose a framework for better quantitative assessment of summary\nquality which is closer to human judgment than a single measure, say F1. We\nreport the performance of a few representative techniques of video\nsummarization on VISIOCITY assessed using various measures and bring out the\nlimitation of the techniques and/or the assessment mechanism in modeling human\njudgment and demonstrate the effectiveness of our evaluation framework in doing\nso.\n