2024/04/29 by Xiangyu Liang, Wenlin Zhuang, Liang, Xiangyu +11 · 1 citation
Computer Science · Mathematics · Psychology · #Animation #Artificial Intelligence (cs.AI) #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer animation #Computer facial animation #Computer graphics (images) #Computer science #Correlation #FOS: Computer and information sciences #Face recognition and analysis #Facial expression #Mathematics #Psychology #Speech recognition
paper · pdf · doi:10.48550/arxiv.2404.18604
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/04/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Speech-driven 3D facial animation technology has been developed for years, but its practical application still lacks expectations. The main challenges lie in data limitations, lip alignment, and the naturalness of facial expressions. Although lip alignment has seen many related studies, existing methods struggle to synthesize natural and realistic expressions, resulting in a mechanical and stiff appearance of facial animations. Even with some research extracting emotional features from speech, the randomness of facial movements limits the effective expression of emotions. To address this issue, this paper proposes a method called CSTalk (Correlation Supervised) that models the correlations among different regions of facial movements and supervises the training of the generative model to generate realistic expressions that conform to human facial motion patterns. To generate more intricate animations, we employ a rich set of control parameters based on the metahuman character model and capture a dataset for five different emotions. We train a generative network using an autoencoder structure and input an emotion embedding vector to achieve the generation of user-control expressions. Experimental results demonstrate that our method outperforms existing state-of-the-art methods.