vix.ing · top · new · best · stats

Toward Multimodal Modeling of Emotional Expressiveness

2020/08/31 by Victoria Lin, Lin, Victoria, Jeffrey M. Girard +6
Computer Science · Mathematics · Neuroscience · Psychology · #Action (physics) #Applications (stat.AP) #Artificial intelligence #Cognitive psychology #Computer science #Emotion and Mood Recognition #Emotional expression #FOS: Computer and information sciences #Face Recognition and Perception #Human-Computer Interaction (cs.HC) #Modalities #Modality (human–computer interaction) #Natural language processing #Psychology #Sentiment Analysis and Opinion Mining #cs.HC #stat.AP

paper · pdf · doi:10.48550/arxiv.2009.00001

published in arXiv (Cornell University) (Cornell University) · V. Lin and J.M. Girard contributed equally to this research. This paper was accepted to ICMI 2020

arxiv created 2020/08/31 · openalex publication_date 2020/08/31 · arxiv updated 2020/09/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

Emotional expressiveness captures the extent to which a person tends to outwardly display their emotions through behavior. Due to the close relationship between emotional expressiveness and behavioral health, as well as the crucial role that it plays in social interaction, the ability to automatically predict emotional expressiveness stands to spur advances in science, medicine, and industry. In this paper, we explore three related research questions. First, how well can emotional expressiveness be predicted from visual, linguistic, and multimodal behavioral signals? Second, which behavioral modalities are uniquely important to the prediction of emotional expressiveness? Third, which behavioral signals are reliably related to emotional expressiveness? To answer these questions, we add highly reliable transcripts and human ratings of perceived emotional expressiveness to an existing video database and use this data to train, validate, and test predictive models. Our best model shows promising predictive performance on this dataset (RMSE=0.65, R2=0.45, r=0.74). Multimodal models tend to perform best overall, and models trained on the linguistic modality tend to outperform models trained on the visual modality. Finally, examination of our interpretable models' coefficients reveals a number of visual and linguistic behavioral signals--such as facial action unit intensity, overall word count, and use of words related to social processes--that reliably predict emotional expressiveness.

Related