2016/07/28 by Stefano Alletto, Giuseppe Serra, Alletto, Stefano +3
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics and Sensor-Based Localization
paper · pdf · doi:10.48550/arxiv.1607.08434
openalex publication_date 2016/07/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
With the spread of wearable devices and head mounted cameras, a wide range of\napplication requiring precise user localization is now possible. In this paper\nwe propose to treat the problem of obtaining the user position with respect to\na known environment as a video registration problem. Video registration, i.e.\nthe task of aligning an input video sequence to a pre-built 3D model, relies on\na matching process of local keypoints extracted on the query sequence to a 3D\npoint cloud. The overall registration performance is strictly tied to the\nactual quality of this 2D-3D matching, and can degrade if environmental\nconditions such as steep changes in lighting like the ones between day and\nnight occur. To effectively register an egocentric video sequence under these\nconditions, we propose to tackle the source of the problem: the matching\nprocess. To overcome the shortcomings of standard matching techniques, we\nintroduce a novel embedding space that allows us to obtain robust matches by\njointly taking into account local descriptors, their spatial arrangement and\ntheir temporal robustness. The proposal is evaluated using unconstrained\negocentric video sequences both in terms of matching quality and resulting\nregistration performance using different 3D models of historical landmarks. The\nresults show that the proposed method can outperform state of the art\nregistration algorithms, in particular when dealing with the challenges of\nnight and day sequences.\n