vix.ing · top · new · best · stats · spec

Historical Document Processing: Historical Document Processing: A Survey of Techniques, Tools, and Trends

2020/02/14 by James Philips, Philips, James P., Nasseh Tabrizi +1
Arts and Humanities · Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Digital Humanities and Scholarship #Digital and Traditional Archives Management #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image Processing and 3D Reconstruction

paper · pdf · doi:10.48550/arxiv.2002.06300

openalex publication_date 2020/02/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Historical Document Processing is the process of digitizing written material from the past for future use by historians and other scholars. It incorporates algorithms and software tools from various subfields of computer science, including computer vision, document analysis and recognition, natural language processing, and machine learning, to convert images of ancient manuscripts, letters, diaries, and early printed texts automatically into a digital format usable in data mining and information retrieval systems. Within the past twenty years, as libraries, museums, and other cultural heritage institutions have scanned an increasing volume of their historical document archives, the need to transcribe the full text from these collections has become acute. Since Historical Document Processing encompasses multiple sub-domains of computer science, knowledge relevant to its purpose is scattered across numerous journals and conference proceedings. This paper surveys the major phases of, standard algorithms, tools, and datasets in the field of Historical Document Processing, discusses the results of a literature review, and finally suggests directions for further research.

Related