2021/03/30 by Yutong Zheng, Yu–Kai Huang, Zheng, Yutong +8 · 2 citations
Computer Science · Mathematics · #Advanced Image Processing Techniques #Artificial intelligence #Cognitive science #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Extrapolation #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Intuition #Mathematics #Natural language processing #Programming language #Representation (politics) #Semantics (computer science) #cs.CV
paper · pdf · doi:10.48550/arxiv.2103.16605
published in arXiv (Cornell University) (Cornell University) · Accepted in IEEE Conference on Computer Vision and Pattern Recognition 2021 (CVPR2021)
arxiv created 2021/03/30 · openalex publication_date 2021/03/30 · arxiv updated 2021/04/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a method to disentangle linear-encoded facial semantics from StyleGAN without external supervision. The method derives from linear regression and sparse representation learning concepts to make the disentangled latent representations easily interpreted as well. We start by coupling StyleGAN with a stabilized 3D deformable facial reconstruction method to decompose single-view GAN generations into multiple semantics. Latent representations are then extracted to capture interpretable facial semantics. In this work, we make it possible to get rid of labels for disentangling meaningful facial semantics. Also, we demonstrate that the guided extrapolation along the disentangled representations can help with data augmentation, which sheds light on handling unbalanced data. Finally, we provide an analysis of our learned localized facial representations and illustrate that the semantic information is encoded, which surprisingly complies with human intuition. The overall unsupervised design brings more flexibility to representation learning in the wild.