2016/11/16 by Kimmo Kettunen, Kettunen, Kimmo
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Mathematics, Computing, and Information Processing #Natural Language Processing Techniques #Web Data Mining and Analysis
paper · pdf · doi:10.48550/arxiv.1611.05239
openalex publication_date 2016/11/16 · openalex created_date 2022/10/06 · openalex updated_date 2026/07/28
The National Library of Finland has digitized the historical newspapers\npublished in Finland between 1771 and 1910. This collection contains\napproximately 1.95 million pages in Finnish and Swedish. Finnish part of the\ncollection consists of about 2.40 billion words. The National Library's Digital\nCollections are offered via the digi.kansalliskirjasto.fi web service, also\nknown as Digi. Part of the newspaper material (from 1771 to 1874) is also\navailable freely downloadable in The Language Bank of Finland provided by the\nFINCLARIN consortium. The collection can also be accessed through the Korp\nenvironment that has been developed by Spr aakbanken at the University of\nGothenburg and extended by FINCLARIN team at the University of Helsinki to\nprovide concordances of text resources. A Cranfield style information retrieval\ntest collection has also been produced out of a small part of the Digi\nnewspaper material at the University of Tampere.\n Quality of OCRed collections is an important topic in digital humanities, as\nit affects general usability and searchability of collections. There is no\nsingle available method to assess quality of large collections, but different\nmethods can be used to approximate quality. This paper discusses different\ncorpus analysis style methods to approximate overall lexical quality of the\nFinnish part of the Digi collection. Methods include usage of parallel samples\nand word error rates, usage of morphological analyzers, frequency analysis of\nwords and comparisons to comparable edited lexical data. Our aim in the quality\nanalysis is twofold: firstly to analyze the present state of the lexical data\nand secondly, to establish a set of assessment methods that build up a compact\nprocedure for quality assessment after e.g. new OCRing or post correction of\nthe material. In the discussion part of the paper we shall synthesize results\nof our different analyses.\n