vix.ing · top · new · best · stats

Imitator: Personalized Speech-driven 3D Facial Animation

2022/12/30 by Balamurugan Thambiraja, Thambiraja, Balamurugan, Ikhsanul Habibie +9 · 28 citations
Computer Science · Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Facial Nerve Paralysis Treatment and Research #Speech and Audio Processing #cs.CV

paper · pdf · doi:10.48550/arxiv.2301.00023

https://youtu.be/JhXTdjiUCUw

arxiv created 2022/12/30 · openalex publication_date 2022/12/30 · arxiv updated 2023/01/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input audio without considering the identity-specific speaking style and facial idiosyncrasies of the target actor, thus, resulting in unrealistic and inaccurate lip movements. To address this, we present Imitator, a speech-driven facial expression synthesis method, which learns identity-specific details from a short input video and produces novel facial expressions matching the identity-specific speaking style and facial idiosyncrasies of the target actor. Specifically, we train a style-agnostic transformer on a large facial expression dataset which we use as a prior for audio-driven facial expressions. Based on this prior, we optimize for identity-specific speaking style based on a short reference video. To train the prior, we introduce a novel loss function based on detected bilabial consonants to ensure plausible lip closures and consequently improve the realism of the generated expressions. Through detailed experiments and a user study, we show that our approach produces temporally coherent facial expressions from input audio while preserving the speaking style of the target actors.

Cited by

Related