2020/12/04 by Parshwa Shah, Shah, Parshwa, Arpit Garg +3
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.2012.02408
openalex publication_date 2020/12/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
A person is usually characterized by descriptors like age, gender, height,\ncloth type, pattern, color, etc. Such descriptors are known as attributes\nand/or soft-biometrics. They link the semantic gap between a person's\ndescription and retrieval in video surveillance. Retrieving a specific person\nwith the query of semantic description has an important application in video\nsurveillance. Using computer vision to fully automate the person retrieval task\nhas been gathering interest within the research community. However, the\nCurrent, trend mainly focuses on retrieving persons with image-based queries,\nwhich have major limitations for practical usage. Instead of using an image\nquery, in this paper, we study the problem of person retrieval in video\nsurveillance with a semantic description. To solve this problem, we develop a\ndeep learning-based cascade filtering approach (PeR-ViS), which uses Mask R-CNN\n[14] (person detection and instance segmentation) and DenseNet-161 [16]\n(soft-biometric classification). On the standard person retrieval dataset of\nSoftBioSearch [6], we achieve 0.566 Average IoU and 0.792 %w IoU > 0.4,\nsurpassing the current state-of-the-art by a large margin. We hope our simple,\nreproducible, and effective approach will help ease future research in the\ndomain of person retrieval in video surveillance. The source code and\npretrained weights available at https://parshwa1999.github.io/PeR-ViS/.\n