2020/08/23 by Jingying Liu, Binyuan Hui, Liu, Jingying +14 · 3 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Graphics (cs.GR) #Human Motion and Animation #cs.GR
paper · pdf · doi:10.48550/arxiv.2008.10004
arxiv created 2020/08/23 · openalex publication_date 2020/08/23 · arxiv updated 2020/08/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Realistic speech-driven 3D facial animation is a challenging problem due to the complex relationship between speech and face. In this paper, we propose a deep architecture, called Geometry-guided Dense Perspective Network (GDPnet), to achieve speaker-independent realistic 3D facial animation. The encoder is designed with dense connections to strengthen feature propagation and encourage the re-use of audio features, and the decoder is integrated with an attention mechanism to adaptively recalibrate point-wise feature responses by explicitly modeling interdependencies between different neuron units. We also introduce a non-linear face reconstruction representation as a guidance of latent space to obtain more accurate deformation, which helps solve the geometry-related deformation and is good for generalization across subjects. Huber and HSIC (Hilbert-Schmidt Independence Criterion) constraints are adopted to promote the robustness of our model and to better exploit the non-linear and high-order correlations. Experimental results on the public dataset and real scanned dataset validate the superiority of our proposed GDPnet compared with state-of-the-art model.