2018/02/21 by Georgios Georgakis, Srikrishna Karanam, Georgakis, Georgios +7 · 1 citation
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Robotics and Sensor-Based Localization
paper · pdf · doi:10.48550/arxiv.1802.07869
openalex publication_date 2018/02/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Finding correspondences between images or 3D scans is at the heart of many\ncomputer vision and image retrieval applications and is often enabled by\nmatching local keypoint descriptors. Various learning approaches have been\napplied in the past to different stages of the matching pipeline, considering\ndetector, descriptor, or metric learning objectives. These objectives were\ntypically addressed separately and most previous work has focused on image\ndata. This paper proposes an end-to-end learning framework for keypoint\ndetection and its representation (descriptor) for 3D depth maps or 3D scans,\nwhere the two can be jointly optimized towards task-specific objectives without\na need for separate annotations. We employ a Siamese architecture augmented by\na sampling layer and a novel score loss function which in turn affects the\nselection of region proposals. The positive and negative examples are obtained\nautomatically by sampling corresponding region proposals based on their\nconsistency with known 3D pose labels. Matching experiments with depth data on\nmultiple benchmark datasets demonstrate the efficacy of the proposed approach,\nshowing significant improvements over state-of-the-art methods.\n