2016/11/24 by Amlan Kar, Kar, Amlan, Nishant Rai +5 · 1 citation
Computer Science · Medicine · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #Diabetic Foot Ulcer Assessment and Management #FOS: Computer and information sciences #Human Pose and Action Recognition
paper · pdf · doi:10.48550/arxiv.1611.08240
openalex publication_date 2016/11/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a novel method for temporally pooling frames in a video for the\ntask of human action recognition. The method is motivated by the observation\nthat there are only a small number of frames which, together, contain\nsufficient information to discriminate an action class present in a video, from\nthe rest. The proposed method learns to pool such discriminative and\ninformative frames, while discarding a majority of the non-informative frames\nin a single temporal scan of the video. Our algorithm does so by continuously\npredicting the discriminative importance of each video frame and subsequently\npooling them in a deep learning framework. We show the effectiveness of our\nproposed pooling method on standard benchmarks where it consistently improves\non baseline pooling methods, with both RGB and optical flow based Convolutional\nnetworks. Further, in combination with complementary video representations, we\nshow results that are competitive with respect to the state-of-the-art results\non two challenging and publicly available benchmark datasets.\n