2023/08/24 by Dimitrios Daskalakis, Daskalakis, Dimitrios, Nikolaos Gkalelis +3
Computer Science · #Advanced Graph Neural Networks #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimedia (cs.MM) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2308.12673
openalex publication_date 2023/08/24 · openalex created_date 2023/08/26 · openalex updated_date 2026/07/28
In this paper, we introduce Masked Feature Modelling (MFM), a novel approach for the unsupervised pre-training of a Graph Attention Network (GAT) block. MFM utilizes a pretrained Visual Tokenizer to reconstruct masked features of objects within a video, leveraging the MiniKinetics dataset. We then incorporate the pre-trained GAT block into a state-of-the-art bottom-up supervised video-event recognition architecture, ViGAT, to improve the model's starting point and overall accuracy. Experimental evaluations on the YLI-MED dataset demonstrate the effectiveness of MFM in improving event recognition performance.