2026/07/13 by Nemanja Rašajski, Konstantinos Makantasis, Antonios Liapis +1
#cs.CV
The 11th Affective Behaviour Analysis in-the-wild Competition includes the Multi-Task Learning Challenge, where participants develop a unified framework for Valence-Arousal Estimation, Expression Recognition, and Action Unit Detection. The challenge lies in learning emotion-related representations that generalize across subjects while remaining robust to spurious factors such as identity, illumination, pose, and demographic variation. To aggregate features extracted by a pre-trained backbone into a compact representation for prediction, attention mechanisms selectively weight the most informative facial regions. However, these attention weights can still capture dataset-specific correlations rather than genuine affective cues. To address this limitation, we propose an attention pooling framework that combines causal supervision with cross-covariance regularization of attention components, encouraging subject-invariant attention and non-redundant representations that improve generalization. Our method achieves CCCVA=0.5123 for VA estimation on the official validation set, together with FEX=0.3116 and FAU=0.3974 for expression recognition and action unit detection, respectively, resulting in an overall P score (the sum of the individual task metrics) of 1.2214.