vix.ing · top · new · best · stats · spec

Infrared and 3D Skeleton Feature Fusion for RGB-D Action Recognition

2020/01/01 by Alban Main de Boissiere, Rita Noumeir
Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Artificial intelligence #Computer science #Computer vision #Convolutional neural network #Feature (linguistics) #Feature extraction #Hand Gesture Recognition Systems #Human Pose and Action Recognition #Modular design #Pattern recognition (psychology) #RGB color model #Skeleton (computer programming) #cs.CV #cs.LG #eess.IV

paper · pdf · doi:10.1109/access.2020.3023599

published as IEEE Access, vol. 8, pp. 168297-168308, 2020 · 11 pages, 5 figures, submitted to IEEE Access

openalex publication_date 2020/01/01 · arxiv created 2020/02/28 · arxiv updated 2021/01/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

For skeleton-based action recognition from depth cameras, distinguishing object-related actions with similar motions is a difficult task. The other available video streams (RGB, infrared, depth) may provide additional clues, given an appropriate feature fusion strategy. We propose a modular network combining skeleton and infrared data. A pre-trained 2D convolutional neural network (CNN) is used as a pose module to extract features from skeleton data. A pre-trained 3D CNN is used as an infrared module to extract visual features from videos. Both feature vectors are then fused and exploited jointly using a multilayer perceptron (MLP). The 2D skeleton coordinates are used to crop a region of interest around the subjects for the infrared videos. Infrared is favored over RGB, as it is less affected by illumination conditions and usable in the dark. We are the first to combine infrared and skeleton data. We evaluate our method on the NTU RGB+D dataset, the largest dataset for human action recognition from depth cameras. We perform extensive ablation studies. In particular, we show the strong contributions of our cropping strategy and pre-training on action classification accuracy. We also test various feature fusion schemes. Feature sum on an element-wise level yields the best results. Our method achieves state-of-the-art performances on NTU RBG+D.

Citations