vix.ing · top · new · best · stats · spec

AdaFuse: Adaptive Temporal Fusion Network for Efficient Action\n Recognition

2021/02/10 by Yue Meng, Meng, Yue, Rameswar Panda +13 · 3 citations
Computer Science · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.2102.05775

openalex publication_date 2021/02/10 · openalex created_date 2021/02/15 · openalex updated_date 2026/07/28

Abstract

Temporal modelling is the key for efficient video action recognition. While\nunderstanding temporal information can improve recognition accuracy for dynamic\nactions, removing temporal redundancy and reusing past features can\nsignificantly save computation leading to efficient action recognition. In this\npaper, we introduce an adaptive temporal fusion network, called AdaFuse, that\ndynamically fuses channels from current and past feature maps for strong\ntemporal modelling. Specifically, the necessary information from the historical\nconvolution feature maps is fused with current pruned feature maps with the\ngoal of improving both recognition accuracy and efficiency. In addition, we use\na skipping operation to further reduce the computation cost of action\nrecognition. Extensive experiments on Something V1 & V2, Jester and\nMini-Kinetics show that our approach can achieve about 40% computation savings\nwith comparable accuracy to state-of-the-art methods. The project page can be\nfound at https://mengyuest.github.io/AdaFuse/\n

Citations

Cited by

Related