2018/05/22 by Ziming Zhang, Rongmei Lin, Zhang, Ziming +3
Computer Science · Engineering · Mathematics · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Robotics and Sensor-Based Localization #cs.CV #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1805.08808
arxiv created 2018/05/22 · openalex publication_date 2018/05/22 · arxiv updated 2018/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper we propose novel Deformable Part Networks (DPNs) to learn \em pose-invariant representations for 2D object recognition. In contrast to the state-of-the-art pose-aware networks such as CapsNet \citesabour2017dynamic and STN \citejaderberg2015spatial, DPNs can be naturally \em interpreted as an efficient solver for a challenging detection problem, namely Localized Deformable Part Models (LDPMs) where localization is introduced to DPMs as another latent variable for searching for the best poses of objects over all pixels and (predefined) scales. In particular we construct DPNs as sequences of such LDPM units to model the semantic and spatial relations among the deformable parts as hierarchical composition and spatial parsing trees. Empirically our 17-layer DPN can outperform both CapsNets and STNs significantly on affNIST \citesabour2017dynamic, for instance, by 19.19% and 12.75%, respectively, with better generalization and better tolerance to affine transformations.