2019/12/08 by Zhenyi Liu, Trisha Lian, Liu, Zhenyi +6 · 4 citations
Computer Science · Mathematics · #Advanced Neural Network Applications #Advanced Vision and Imaging #Artificial intelligence #Artificial neural network #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Convolutional neural network #FOS: Computer and information sciences #Generalization #Image sensor #Mathematics #Multispectral image #Pattern recognition (psychology) #Pixel #Video Surveillance and Tracking Methods #cs.CV
paper · pdf · doi:10.48550/arxiv.1912.03604
published in arXiv (Cornell University) (Cornell University) · 11 pages, 11 figures, in preparation for submission
arxiv created 2019/12/08 · openalex publication_date 2019/12/08 · arxiv updated 2019/12/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We quantify the generalization of a convolutional neural network (CNN) trained to identify cars. First, we perform a series of experiments to train the network using one image dataset - either synthetic or from a camera - and then test on a different image dataset. We show that generalization between images obtained with different cameras is roughly the same as generalization between images from a camera and ray-traced multispectral synthetic images. Second, we use ISETAuto, a soft prototyping tool that creates ray-traced multispectral simulations of camera images, to simulate sensor images with a range of pixel sizes, color filters, acquisition and post-acquisition processing. These experiments reveal how variations in specific camera parameters and image processing operations impact CNN generalization. We find that (a) pixel size impacts generalization, (b) demosaicking substantially impacts performance and generalization for shallow (8-bit) bit-depths but not deeper ones (10-bit), and (c) the network performs well using raw (not demosaicked) sensor data for 10-bit pixels.