vix.ing · top · new · best · stats · spec

LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from\n Video using Pose and Lighting Normalization

2021/06/08 by Avisek Lahiri, Vivek Kwatra, Lahiri, Avisek +7 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing

paper · pdf · doi:10.48550/arxiv.2106.04185

openalex publication_date 2021/06/08 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

In this paper, we present a video-based learning framework for animating\npersonalized 3D talking faces from audio. We introduce two training-time data\nnormalizations that significantly improve data sample efficiency. First, we\nisolate and represent faces in a normalized space that decouples 3D geometry,\nhead pose, and texture. This decomposes the prediction problem into regressions\nover the 3D face shape and the corresponding 2D texture atlas. Second, we\nleverage facial symmetry and approximate albedo constancy of skin to isolate\nand remove spatio-temporal lighting variations. Together, these normalizations\nallow simple networks to generate high fidelity lip-sync videos under novel\nambient illumination while training with just a single speaker-specific video.\nFurther, to stabilize temporal dynamics, we introduce an auto-regressive\napproach that conditions the model on its previous visual state. Human ratings\nand objective metrics demonstrate that our method outperforms contemporary\nstate-of-the-art audio-driven video reenactment benchmarks in terms of realism,\nlip-sync and visual quality scores. We illustrate several applications enabled\nby our framework.\n

Cited by

Related