2019/11/19 by Barak Battash, Battash, Barak, Haim Barad +5
Computer Science · #Human Pose and Action Recognition #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications
paper · pdf · doi:10.48550/arxiv.1911.08206
Video understanding usually requires expensive computation that prohibits its\ndeployment, yet videos contain significant spatiotemporal redundancy that can\nbe exploited. In particular, operating directly on the motion vectors and\nresiduals in the compressed video domain can significantly accelerate the\ncompute, by not using the raw videos which demand colossal storage capacity.\nExisting methods approach this task as a multiple modalities problem. In this\npaper we are approaching the task in a completely different way; we are looking\nat the data from the compressed stream as a one unit clip and propose that the\nresidual frames can replace the original RGB frames from the raw domain.\nFurthermore, we are using teacher-student method to aid the network in the\ncompressed domain to mimic the teacher network in the raw domain. We show\nexperiments on three leading datasets (HMDB51, UCF1, and Kinetics) that\napproach state-of-the-art accuracy on raw video data by using compressed data.\nOur model MFCD-Net outperforms prior methods in the compressed domain and more\nimportantly, our model has 11X fewer parameters and 3X fewer Flops,\ndramatically improving the efficiency of video recognition inference. This\napproach enables applying neural networks exclusively in the compressed domain\nwithout compromising accuracy while accelerating performance.\n