2010/03/30 by Sandip Rakshit, Rakshit, Sandip, Subhadip Basu +3
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image Processing and 3D Reconstruction #Vehicle License Plate Recognition #cs.CV
paper · pdf · doi:10.48550/arxiv.1003.5893
Proc. Int. Conf. on Information Technology and Business Intelligence (2009) 117-125
arxiv created 2010/03/30 · openalex publication_date 2010/03/30 · arxiv updated 2010/03/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02
Objective of the current work is to develop an Optical Character Recognition (OCR) engine for information Just In Time (iJIT) system that can be used for recognition of handwritten textual annotations of lower case Roman script. Tesseract open source OCR engine under Apache License 2.0 is used to develop user-specific handwriting recognition models, viz., the language sets, for the said system, where each user is identified by a unique identification tag associated with the digital pen. To generate the language set for any user, Tesseract is trained with labeled handwritten data samples of isolated and free-flow texts of Roman script, collected exclusively from that user. The designed system is tested on five different language sets with free- flow handwritten annotations as test samples. The system could successfully segment and subsequently recognize 87.92%, 81.53%, 92.88%, 86.75% and 90.80% handwritten characters in the test samples of five different users.