vix.ing · top · new · best · stats · spec

Where, What, Whether: Multi-modal Learning Meets Pedestrian Detection

2020/12/20 by Yan Luo, Chongyang Zhang, Luo, Yan +7 · 1 citation
Computer Science · #Advanced Neural Network Applications #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Surveillance and Tracking Methods

paper · pdf · doi:10.48550/arxiv.2012.10880

openalex publication_date 2020/12/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Pedestrian detection benefits greatly from deep convolutional neural networks (CNNs). However, it is inherently hard for CNNs to handle situations in the presence of occlusion and scale variation. In this paper, we propose W3Net, which attempts to address above challenges by decomposing the pedestrian detection task into \textbfWhere, \textbfWhat and \textbfWhether problem directing against pedestrian localization, scale prediction and classification correspondingly. Specifically, for a pedestrian instance, we formulate its feature by three steps. i) We generate a bird view map, which is naturally free from occlusion issues, and scan all points on it to look for suitable locations for each pedestrian instance. ii) Instead of utilizing pre-fixed anchors, we model the interdependency between depth and scale aiming at generating depth-guided scales at different locations for better matching instances of different sizes. iii) We learn a latent vector shared by both visual and corpus space, by which false positives with similar vertical structure but lacking human partial features would be filtered out. We achieve state-of-the-art results on widely used datasets (Citypersons and Caltech). In particular. when evaluating on heavy occlusion subset, our results reduce MR-2 from 49.3% to 18.7% on Citypersons, and from 45.18% to 28.33% on Caltech.

Citations

Cited by

Related