2018/04/14 by Zhi-Qi Cheng, Cheng, Zhi-Qi, Xiao Wu +6 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #cs.CV
paper · pdf · doi:10.48550/arxiv.1804.05287
IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2017
openalex publication_date 2018/04/14 · arxiv created 2018/12/04 · arxiv updated 2018/12/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In recent years, both online retail and video hosting service are exponentially growing. In this paper, we explore a new cross-domain task, Video2Shop, targeting for matching clothes appeared in videos to the exact same items in online shops. A novel deep neural network, called AsymNet, is proposed to explore this problem. For the image side, well-established methods are used to detect and extract features for clothing patches with arbitrary sizes. For the video side, deep visual features are extracted from detected object regions in each frame, and further fed into a Long Short-Term Memory (LSTM) framework for sequence modeling, which captures the temporal dynamics in videos. To conduct exact matching between videos and online shopping images, LSTM hidden states, representing the video, and image features, which represent static object images, are jointly modeled under the similarity network with reconfigurable deep tree structure. Moreover, an approximate training method is proposed to achieve the efficiency when training. Extensive experiments conducted on a large cross-domain dataset have demonstrated the effectiveness and efficiency of the proposed AsymNet, which outperforms the state-of-the-art methods.