vix.ing · top · new · best · stats · spec

Augmented Co-Speech Gesture Generation: Including Form and Meaning Features to Guide Learning-Based Gesture Synthesis

2023/07/13 by Hendric Voß, Voß, Hendric, Stefan Kopp +1 · 3 citations
Computer Science · #FOS: Computer and information sciences #Graphics (cs.GR) #Hand Gesture Recognition Systems #Human-Computer Interaction (cs.HC) #Multimodal Machine Learning Applications #Speech and dialogue systems

paper · pdf · doi:10.48550/arxiv.2307.09597

openalex publication_date 2023/07/13 · openalex created_date 2023/07/21 · openalex updated_date 2026/07/28

Abstract

Due to their significance in human communication, the automatic generation of co-speech gestures in artificial embodied agents has received a lot of attention. Although modern deep learning approaches can generate realistic-looking conversational gestures from spoken language, they often lack the ability to convey meaningful information and generate contextually appropriate gestures. This paper presents an augmented approach to the generation of co-speech gestures that additionally takes into account given form and meaning features for the gestures. Our framework effectively acquires this information from a small corpus with rich semantic annotations and a larger corpus without such information. We provide an analysis of the effects of distinctive feature targets and we report on a human rater evaluation study demonstrating that our framework achieves semantic coherence and person perception on the same level as human ground truth behavior. We make our data pipeline and the generation framework publicly available.

Cited by

Related