2016/11/27 by Aaron Chadha, Chadha, Aaron, Yiannis Andreopoulos +1 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1611.08906
openalex publication_date 2016/11/27 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28
We investigate the problem of image retrieval based on visual queries when\nthe latter comprise arbitrary regions-of-interest (ROI) rather than entire\nimages. Our proposal is a compact image descriptor that combines the\nstate-of-the-art in content-based descriptor extraction with a multi-level,\nVoronoi-based spatial partitioning of each dataset image. The proposed\nmulti-level Voronoi-based encoding uses a spatial hierarchical K-means over\ninterest-point locations, and computes a content-based descriptor over each\ncell. In order to reduce the matching complexity with minimal or no sacrifice\nin retrieval performance: (i) we utilize the tree structure of the spatial\nhierarchical K-means to perform a top-to-bottom pruning for local similarity\nmaxima; (ii) we propose a new image similarity score that combines relevant\ninformation from all partition levels into a single measure for similarity;\n(iii) we combine our proposal with a novel and efficient approach for optimal\nbit allocation within quantized descriptor representations. By deriving both a\nVoronoi-based VLAD descriptor (termed as Fast-VVLAD) and a Voronoi-based deep\nconvolutional neural network (CNN) descriptor (termed as Fast-VDCNN), we\ndemonstrate that our Voronoi-based framework is agnostic to the descriptor\nbasis, and can easily be slotted into existing frameworks. Via a range of ROI\nqueries in two standard datasets, it is shown that the Voronoi-based\ndescriptors achieve comparable or higher mean Average Precision against\nconventional grid-based spatial search, while offering more than two-fold\nreduction in complexity. Finally, beyond ROI queries, we show that Voronoi\npartitioning improves the geometric invariance of compact CNN descriptors,\nthereby resulting in competitive performance to the current state-of-the-art on\nwhole image retrieval.\n