2025/05/09 by Maan Alhazmi, Alhazmi, Maan, Abdulrahman Altahhan +1
Computer Science · Neuroscience · Psychology · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Emotion and Mood Recognition #FOS: Computer and information sciences #Face Recognition and Perception #Face recognition and analysis #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2505.05943
openalex publication_date 2025/05/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
The emergence of ConvNeXt and its variants has reaffirmed the conceptual and structural suitability of CNN-based models for vision tasks, re-establishing them as key players in image classification in general, and in facial expression recognition (FER) in particular. In this paper, we propose a new set of models that build on these advancements by incorporating a new set of attention mechanisms that combines Triplet attention with Squeeze-and-Excitation (TripSE) in four different variants. We demonstrate the effectiveness of these variants by applying them to the ResNet18, DenseNet and ConvNext architectures to validate their versatility and impact. Our study shows that incorporating a TripSE block in these CNN models boosts their performances, particularly for the ConvNeXt architecture, indicating its utility. We evaluate the proposed mechanisms and associated models across four datasets, namely CIFAR100, ImageNet, FER2013 and AffectNet datasets, where ConvNext with TripSE achieves state-of-the-art results with an accuracy of 78.27% on the popular FER2013 dataset, a new feat for this dataset.