vix.ing · top · new · best · stats · spec

Investigation on Combining 3D Convolution of Image Data and Optical Flow\n to Generate Temporal Action Proposals

2019/03/11 by Patrick Schlosser, Schlosser, Patrick, David Münch +3
Computer Science · #Advanced Vision and Imaging #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition

paper · pdf · doi:10.48550/arxiv.1903.04176

openalex publication_date 2019/03/11 · openalex created_date 2022/07/29 · openalex updated_date 2026/07/28

Abstract

In this paper, several variants of two-stream architectures for temporal\naction proposal generation in long, untrimmed videos are presented. Inspired by\nthe recent advances in the field of human action recognition utilizing 3D\nconvolutions in combination with two-stream networks and based on the\nSingle-Stream Temporal Action Proposals (SST) architecture, four different\ntwo-stream architectures utilizing sequences of images on one stream and\nsequences of images of optical flow on the other stream are subsequently\ninvestigated. The four architectures fuse the two separate streams at different\ndepths in the model; for each of them, a broad range of parameters is\ninvestigated systematically as well as an optimal parametrization is\nempirically determined. The experiments on the THUMOS'14 dataset show that all\nfour two-stream architectures are able to outperform the original single-stream\nSST and achieve state of the art results. Additional experiments revealed that\nthe improvements are not restricted to a single method of calculating optical\nflow by exchanging the formerly used method of Brox with FlowNet2 and still\nachieving improvements.\n

Related