2015/09/25 by Hirokatsu Kataoka, Kataoka, Hirokatsu, Kenji Iwata +3
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Video Surveillance and Tracking Methods #cs.AI #cs.CV #cs.MM
paper · pdf · doi:10.48550/arxiv.1509.07627
5 pages, 3 figures
arxiv created 2015/09/25 · openalex publication_date 2015/09/25 · arxiv updated 2015/09/28 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
In this paper, we evaluate convolutional neural network (CNN) features using the AlexNet architecture and very deep convolutional network (VGGNet) architecture. To date, most CNN researchers have employed the last layers before output, which were extracted from the fully connected feature layers. However, since it is unlikely that feature representation effectiveness is dependent on the problem, this study evaluates additional convolutional layers that are adjacent to fully connected layers, in addition to executing simple tuning for feature concatenation (e.g., layer 3 + layer 5 + layer 7) and transformation, using tools such as principal component analysis. In our experiments, we carried out detection and classification tasks using the Caltech 101 and Daimler Pedestrian Benchmark Datasets.