2022/03/25 by Jun-Hwa Kim, Nam‐Ho Kim, Kim, Jun-Hwa +4
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Emotion and Mood Recognition #FOS: Computer and information sciences #Face and Expression Recognition #Speech and Audio Processing #cs.AI #cs.CV
paper · pdf · doi:10.48550/arxiv.2203.13472
arxiv created 2022/03/25 · openalex publication_date 2022/03/25 · arxiv updated 2022/03/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The task of recognizing human facial expressions plays a vital role in various human-related systems, including health care and medical fields. With the recent success of deep learning and the accessibility of a large amount of annotated data, facial expression recognition research has been mature enough to be utilized in real-world scenarios with audio-visual datasets. In this paper, we introduce Swin transformer-based facial expression approach for an in-the-wild audio-visual dataset of the Aff-Wild2 Expression dataset. Specifically, we employ a three-stream network (i.e., Visual stream, Temporal stream, and Audio stream) for the audio-visual videos to fuse the multi-modal information into facial expression recognition. Experimental results on the Aff-Wild2 dataset show the effectiveness of our proposed multi-modal approaches.