2021/01/01 by Xiaofeng Liu, Linghao Jin, Liu, Xiaofeng +9 · 1 citation
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Emotion and Mood Recognition #FOS: Computer and information sciences #Face recognition and analysis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Speech and Audio Processing #cs.AI #cs.CV #cs.LG #cs.MM
paper · pdf · doi:10.48550/arxiv.2101.00317
Accepted as the Oral paper at ICPR 2020 (<4.4%). arXiv admin note: substantial text overlap with arXiv:2010.10637
openalex publication_date 2021/01/01 · arxiv created 2021/01/07 · arxiv updated 2021/01/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper targets to explore the inter-subject variations eliminated facial expression representation in the compressed video domain. Most of the previous methods process the RGB images of a sequence, while the off-the-shelf and valuable expression-related muscle movement already embedded in the compression format. In the up to two orders of magnitude compressed domain, we can explicitly infer the expression from the residual frames and possible to extract identity factors from the I frame with a pre-trained face recognition network. By enforcing the marginal independent of them, the expression feature is expected to be purer for the expression and be robust to identity shifts. We do not need the identity label or multiple expression samples from the same person for identity elimination. Moreover, when the apex frame is annotated in the dataset, the complementary constraint can be further added to regularize the feature-level game. In testing, only the compressed residual frames are required to achieve expression prediction. Our solution can achieve comparable or better performance than the recent decoded image based methods on the typical FER benchmarks with about 3× faster inference with compressed data.