2021/05/17 by Ayantha Randika, Randika, Ayantha, Nilanjan Ray +5
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image Retrieval and Classification Techniques
paper · pdf · doi:10.48550/arxiv.2105.07983
openalex publication_date 2021/05/17 · openalex created_date 2022/09/16 · openalex updated_date 2026/07/28
Optical character recognition (OCR) is a widely used pattern recognition\napplication in numerous domains. There are several feature-rich,\ngeneral-purpose OCR solutions available for consumers, which can provide\nmoderate to excellent accuracy levels. However, accuracy can diminish with\ndifficult and uncommon document domains. Preprocessing of document images can\nbe used to minimize the effect of domain shift. In this paper, a novel approach\nis presented for creating a customized preprocessor for a given OCR engine.\nUnlike the previous OCR agnostic preprocessing techniques, the proposed\napproach approximates the gradient of a particular OCR engine to train a\npreprocessor module. Experiments with two datasets and two OCR engines show\nthat the presented preprocessor is able to improve the accuracy of the OCR up\nto 46% from the baseline by applying pixel-level manipulations to the document\nimage. The implementation of the proposed method and the enhanced public\ndatasets are available for download.\n