2016/08/26 by Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li +1 · 44 citations
Computer Science · #Biometric Identification and Security #Face and Expression Recognition #Face recognition and analysis
paper · doi:10.1109/lsp.2016.2603342
openalex created_date 2016/06/24 · openalex publication_date 2016/08/26 · crossref created 2016/08/26 · crossref issued 2016/10/01 · crossref published 2016/10/01 · crossref published-print 2016/10/01 · crossref deposited 2022/01/12 · crossref indexed 2026/07/30 · openalex updated_date 2026/08/01
Face detection and alignment in unconstrained environment are challenging due to various poses, illuminations, and occlusions. Recent studies show that deep learning approaches can achieve impressive performance on these two tasks. In this letter, we propose a deep cascaded multitask framework that exploits the inherent correlation between detection and alignment to boost up their performance. In particular, our framework leverages a cascaded architecture with three stages of carefully designed deep convolutional networks to predict face and landmark location in a coarse-to-fine manner. In addition, we propose a new online hard sample mining strategy that further improves the performance in practice. Our method achieves superior accuracy over the state-of-the-art techniques on the challenging face detection dataset and benchmark and WIDER FACE benchmarks for face detection, and annotated facial landmarks in the wild benchmark for face alignment, while keeps real-time performance.