2017/07/21 by Rahul Sharma, Tanaya Guha, Sharma, Rahul +3
Arts and Humanities · Computer Science · #FOS: Computer and information sciences #Image and Video Quality Assessment #Multimedia (cs.MM) #Subtitles and Audiovisual Media #Visual Attention and Saliency Detection
paper · pdf · doi:10.48550/arxiv.1707.06830
openalex publication_date 2017/07/21 · openalex created_date 2022/10/05 · openalex updated_date 2026/07/28
Public speaking is an important aspect of human communication and\ninteraction. The majority of computational work on public speaking concentrates\non analyzing the spoken content, and the verbal behavior of the speakers. While\nthe success of public speaking largely depends on the content of the talk, and\nthe verbal behavior, non-verbal (visual) cues, such as gestures and physical\nappearance also play a significant role. This paper investigates the importance\nof visual cues by estimating their contribution towards predicting the\npopularity of a public lecture. For this purpose, we constructed a large\ndatabase of more than 1800 TED talk videos. As a measure of popularity of the\nTED talks, we leverage the corresponding (online) viewers' ratings from\nYouTube. Visual cues related to facial and physical appearance, facial\nexpressions, and pose variations are extracted from the video frames using\nconvolutional neural network (CNN) models. Thereafter, an attention-based long\nshort-term memory (LSTM) network is proposed to predict the video popularity\nfrom the sequence of visual features. The proposed network achieves\nstate-of-the-art prediction accuracy indicating that visual cues alone contain\nhighly predictive information about the popularity of a talk. Furthermore, our\nnetwork learns a human-like attention mechanism, which is particularly useful\nfor interpretability, i.e. how attention varies with time, and across different\nvisual cues by indicating their relative importance.\n