2016/08/25 by César Roberto de Souza, de Souza, César Roberto, Adrien Gaidon +5
Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition
paper · pdf · doi:10.48550/arxiv.1608.07138
openalex publication_date 2016/08/25 · openalex created_date 2022/10/06 · openalex updated_date 2026/07/28
Action recognition in videos is a challenging task due to the complexity of\nthe spatio-temporal patterns to model and the difficulty to acquire and learn\non large quantities of video data. Deep learning, although a breakthrough for\nimage classification and showing promise for videos, has still not clearly\nsuperseded action recognition methods using hand-crafted features, even when\ntraining on massive datasets. In this paper, we introduce hybrid video\nclassification architectures based on carefully designed unsupervised\nrepresentations of hand-crafted spatio-temporal features classified by\nsupervised deep networks. As we show in our experiments on five popular\nbenchmarks for action recognition, our hybrid model combines the best of both\nworlds: it is data efficient (trained on 150 to 10000 short clips) and yet\nimproves significantly on the state of the art, including recent deep models\ntrained on millions of manually labelled images and videos.\n