2017/12/06 by Frédéric Rayar, Rayar, Frédéric, Masanori Goto +3
Computer Science · #Advanced Neural Network Applications #FOS: Computer and information sciences #Face and Expression Recognition #Handwritten Text Recognition Techniques #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.1712.02122
openalex publication_date 2017/12/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we present a study on sample preselection in large training data set for CNN-based classification. To do so, we structure the input data set in a network representation, namely the Relative Neighbourhood Graph, and then extract some vectors of interest. The proposed preselection method is evaluated in the context of handwritten character recognition, by using two data sets, up to several hundred thousands of images. It is shown that the graph-based preselection can reduce the training data set without degrading the recognition accuracy of a non pretrained CNN shallow model.