2019/10/30 by Yudong Jiang, Jiang, Yudong, Kaixu Cui +5
Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia Communication and Technology #Music and Audio Processing #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.1910.13888
openalex publication_date 2019/10/30 · openalex created_date 2020/07/23 · openalex updated_date 2026/07/28
Video summarization aims to extract keyframes/shots from a long video.\nPrevious methods mainly take diversity and representativeness of generated\nsummaries as prior knowledge in algorithm design. In this paper, we formulate\nvideo summarization as a content-based recommender problem, which should\ndistill the most useful content from a long video for users who suffer from\ninformation overload. A scalable deep neural network is proposed on predicting\nif one video segment is a useful segment for users by explicitly modelling both\nsegment and video. Moreover, we accomplish scene and action recognition in\nuntrimmed videos in order to find more correlations among different aspects of\nvideo understanding tasks. Also, our paper will discuss the effect of audio and\nvisual features in summarization task. We also extend our work by data\naugmentation and multi-task learning for preventing the model from early-stage\noverfitting. The final results of our model win the first place in ICCV 2019\nCoView Workshop Challenge Track.\n