2017/11/02 by I‐Hong Jhuo, I-Hong Jhuo, Jun Wang +2
Computer Science · Mathematics · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Data mining #FOS: Computer and information sciences #Hash function #Hash table #Image (mathematics) #Image Retrieval and Classification Techniques #Kernel (algebra) #Kernel method #Locality-sensitive hashing #Machine Learning (cs.LG) #Mathematics #Multimedia (cs.MM) #Nearest neighbor search #Pattern recognition (psychology) #Representation (politics) #Set (abstract data type) #Similarity (geometry) #Support vector machine #Theoretical computer science #Video Surveillance and Tracking Methods #cs.CV #cs.LG #cs.MM
paper · pdf · doi:10.48550/arxiv.1711.00888
published in arXiv (Cornell University) (Cornell University) · 9 pages
openalex publication_date 2017/11/02 · arxiv created 2019/05/29 · arxiv updated 2019/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Visual data, such as an image or a sequence of video frames, is often naturally represented as a point set. In this paper, we consider the fundamental problem of finding a nearest set from a collection of sets, to a query set. This problem has obvious applications in large-scale visual retrieval and recognition, and also in applied fields beyond computer vision. One challenge stands out in solving the problem---set representation and measure of similarity. Particularly, the query set and the sets in dataset collection can have varying cardinalities. The training collection is large enough such that linear scan is impractical. We propose a simple representation scheme that encodes both statistical and structural information of the sets. The derived representations are integrated in a kernel framework for flexible similarity measurement. For the query set process, we adopt a learning-to-hash pipeline that turns the kernel representations into hash bits based on simple learners, using multiple kernel learning. Experiments on two visual retrieval datasets show unambiguously that our set-to-set hashing framework outperforms prior methods that do not take the set-to-set search setting.