vix.ing · top · new · best · stats · spec

Mixed Model OCR Training on Historical Latin Script for Out-of-the-Box\n Recognition and Finetuning

2021/06/15 by Christian Reul, Christoph Wick, Reul, Christian +9 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image Processing and 3D Reconstruction #Mathematics, Computing, and Information Processing

paper · pdf · doi:10.48550/arxiv.2106.07881

openalex publication_date 2021/06/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In order to apply Optical Character Recognition (OCR) to historical printings\nof Latin script fully automatically, we report on our efforts to construct a\nwidely-applicable polyfont recognition model yielding text with a Character\nError Rate (CER) around 2% when applied out-of-the-box. Moreover, we show how\nthis model can be further finetuned to specific classes of printings with\nlittle manual and computational effort. The mixed or polyfont model is trained\non a wide variety of materials, in terms of age (from the 15th to the 19th\ncentury), typography (various types of Fraktur and Antiqua), and languages\n(among others, German, Latin, and French). To optimize the results we combined\nestablished techniques of OCR training like pretraining, data augmentation, and\nvoting. In addition, we used various preprocessing methods to enrich the\ntraining data and obtain more robust models. We also implemented a two-stage\napproach which first trains on all available, considerably unbalanced data and\nthen refines the output by training on a selected more balanced subset.\nEvaluations on 29 previously unseen books resulted in a CER of 1.73%,\noutperforming a widely used standard model with a CER of 2.84% by almost 40%.\nTraining a more specialized model for some unseen Early Modern Latin books\nstarting from our mixed model led to a CER of 1.47%, an improvement of up to\n50% compared to training from scratch and up to 30% compared to training from\nthe aforementioned standard model. Our new mixed model is made openly available\nto the community.\n

Cited by

Related