1996/04/24 by Geunbae Lee, Lee, Geunbae, Jong-Hyeok Lee +4
Computer Science · #Advanced Text Analysis Techniques #Bayesian Modeling and Causal Inference #Computation and Language (cs.CL) #FOS: Computer and information sciences #Text and Document Classification Technologies #cmp-lg #cs.CL
paper · pdf · doi:10.48550/arxiv.cmp-lg/9604011
latex with a4, epsfig style, 21 pages, 11 postscript figures, accepted in pattern recognition journal
arxiv created 1996/04/24 · openalex publication_date 1996/04/24 · arxiv updated 2009/11/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Most of the post-processing methods for character recognition rely on contextual information of character and word-fragment levels. However, due to linguistic characteristics of Korean, such low-level information alone is not sufficient for high-quality character-recognition applications, and we need much higher-level contextual information to improve the recognition results. This paper presents a domain independent post-processing technique that utilizes multi-level morphological, syntactic, and semantic information as well as character-level information. The proposed post-processing system performs three-level processing: candidate character-set selection, candidate eojeol (Korean word) generation through morphological analysis, and final single eojeol-sequence selection by linguistic evaluation. All the required linguistic information and probabilities are automatically acquired from a statistical corpus analysis. Experimental results demonstrate the effectiveness of our method, yielding error correction rate of 80.46%, and improved recognition rate of 95.53% from before-post-processing rate 71.2% for single best-solution selection.