2020/03/16 by Osman Semih Kayhan, Kayhan, Osman Semih, Jan van Gemert +1 · 10 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2003.07064
openalex publication_date 2020/03/16 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
In this paper we challenge the common assumption that convolutional layers in\nmodern CNNs are translation invariant. We show that CNNs can and will exploit\nthe absolute spatial location by learning filters that respond exclusively to\nparticular absolute locations by exploiting image boundary effects. Because\nmodern CNNs filters have a huge receptive field, these boundary effects operate\neven far from the image boundary, allowing the network to exploit absolute\nspatial location all over the image. We give a simple solution to remove\nspatial location encoding which improves translation invariance and thus gives\na stronger visual inductive bias which particularly benefits small data sets.\nWe broadly demonstrate these benefits on several architectures and various\napplications such as image classification, patch matching, and two video\nclassification datasets.\n