2020/02/04 by Huabin Wang, Wang, Huabin, Rui Cheng +7 · 1 citation
Computer Science · Engineering · #Algorithm #Artificial intelligence #Benchmark (surveying) #Bounding overwatch #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Convolutional neural network #Deep learning #Deep neural networks #FOS: Computer and information sciences #FOS: Electrical engineering #Face (sociological concept) #Face and Expression Recognition #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Hourglass #Image and Video Processing (eess.IV) #Initialization #Pattern recognition (psychology) #Residual #Transformer #cs.CV #eess.IV #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2002.01075
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/02/04 · openalex publication_date 2020/02/04 · arxiv updated 2020/02/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
An ability to generalize unconstrained conditions such as severe occlusions and large pose variations remains a challenging goal to achieve in face alignment. In this paper, a multistage model based on deep neural networks is proposed which takes advantage of spatial transformer networks, hourglass networks and exemplar-based shape constraints. First, a spatial transformer - generative adversarial network which consists of convolutional layers and residual units is utilized to solve the initialization issues caused by face detectors, such as rotation and scale variations, to obtain improved face bounding boxes for face alignment. Then, stacked hourglass network is employed to obtain preliminary locations of landmarks as well as their corresponding scores. In addition, an exemplar-based shape dictionary is designed to determine landmarks with low scores based on those with high scores. By incorporating face shape constraints, misaligned landmarks caused by occlusions or cluttered backgrounds can be considerably improved. Extensive experiments based on challenging benchmark datasets are performed to demonstrate the superior performance of the proposed method over other state-of-the-art methods.