2020/11/30 by Alok Singh, Singh, Alok, Thoudam Doren Singh +3
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.2011.14752
openalex publication_date 2020/11/30 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Video description involves the generation of the natural language description\nof actions, events, and objects in the video. There are various applications of\nvideo description by filling the gap between languages and vision for visually\nimpaired people, generating automatic title suggestion based on content,\nbrowsing of the video based on the content and video-guided machine translation\n[86] etc.In the past decade, several works had been done in this field in terms\nof approaches/methods for video description, evaluation metrics,and datasets.\nFor analyzing the progress in the video description task, a comprehensive\nsurvey is needed that covers all the phases of video description approaches\nwith a special focus on recent deep learning approaches. In this work, we\nreport a comprehensive survey on the phases of video description approaches,\nthe dataset for video description, evaluation metrics, open competitions for\nmotivating the research on the video description, open challenges in this\nfield, and future research directions. In this survey, we cover the\nstate-of-the-art approaches proposed for each and every dataset with their pros\nand cons. For the growth of this research domain,the availability of numerous\nbenchmark dataset is a basic need. Further, we categorize all the dataset into\ntwo classes: open domain dataset and domain-specific dataset. From our survey,\nwe observe that the work in this field is in fast-paced development since the\ntask of video description falls in the intersection of computer vision and\nnatural language processing. But still, the work in the video description is\nfar from saturation stage due to various challenges like the redundancy due to\nsimilar frames which affect the quality of visual features, the availability of\ndataset containing more diverse content and availability of an effective\nevaluation metric.\n