2017/07/27 by Christian Bartz, Bartz, Christian, Haojin Yang +3 · 51 citations
Computer Science · Engineering · #Artificial intelligence #Artificial neural network #Benchmark (surveying) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image (mathematics) #Image Processing and 3D Reconstruction #Natural language processing #Pattern recognition (psychology) #Text detection #Text recognition #Transformer #Vehicle License Plate Recognition #cs.CV
paper · pdf · doi:10.48550/arxiv.1707.08831
published in arXiv (Cornell University) (Cornell University) · 9 pages, 6 figures
arxiv created 2017/07/27 · openalex publication_date 2017/07/27 · arxiv updated 2017/07/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Detecting and recognizing text in natural scene images is a challenging, yet not completely solved task. In re- cent years several new systems that try to solve at least one of the two sub-tasks (text detection and text recognition) have been proposed. In this paper we present STN-OCR, a step towards semi-supervised neural networks for scene text recognition, that can be optimized end-to-end. In contrast to most existing works that consist of multiple deep neural networks and several pre-processing steps we propose to use a single deep neural network that learns to detect and recognize text from natural images in a semi-supervised way. STN-OCR is a network that integrates and jointly learns a spatial transformer network, that can learn to detect text regions in an image, and a text recognition network that takes the identified text regions and recognizes their textual content. We investigate how our model behaves on a range of different tasks (detection and recognition of characters, and lines of text). Experimental results on public benchmark datasets show the ability of our model to handle a variety of different tasks, without substantial changes in its overall network structure.