vix.ing · top · new · best · stats · spec

Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video\n Action Recognition

2018/08/22 by Unaiza Ahsan, Ahsan, Unaiza, Rishi Madhok +3 · 2 citations
Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition

paper · pdf · doi:10.48550/arxiv.1808.07507

openalex publication_date 2018/08/22 · openalex created_date 2022/08/04 · openalex updated_date 2026/07/28

Abstract

We propose a self-supervised learning method to jointly reason about spatial\nand temporal context for video recognition. Recent self-supervised approaches\nhave used spatial context [9, 34] as well as temporal coherency [32] but a\ncombination of the two requires extensive preprocessing such as tracking\nobjects through millions of video frames [59] or computing optical flow to\ndetermine frame regions with high motion [30]. We propose to combine spatial\nand temporal context in one self-supervised framework without any heavy\npreprocessing. We divide multiple video frames into grids of patches and train\na network to solve jigsaw puzzles on these patches from multiple frames. So the\nnetwork is trained to correctly identify the position of a patch within a video\nframe as well as the position of a patch over time. We also propose a novel\npermutation strategy that outperforms random permutations while significantly\nreducing computational and memory constraints. We use our trained network for\ntransfer learning tasks such as video activity recognition and demonstrate the\nstrength of our approach on two benchmark video action recognition datasets\nwithout using a single frame from these datasets for unsupervised pretraining\nof our proposed video jigsaw network.\n

Citations

Cited by

Related