vix.ing · top · new · best · stats · spec

Adapting Tesseract for Complex Scripts: An Example for Urdu Nastalique

2014/04/01 by Qurat ul Ain Akram, Sarmad Hussain, Aneta Niazi +2 · 1 citation
Computer Science · Engineering · #Handwritten Text Recognition Techniques #Advanced Image and Video Retrieval Techniques #Vehicle License Plate Recognition #Cursive #Computer science #Urdu #Artificial intelligence #Scripting language #Arabic #Natural language processing #Speech recognition #Linguistics #Programming language

paper · doi:10.1109/das.2014.45

openalex publication_date 2014/04/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29

Abstract

Tesseract engine supports multilingual text recognition. However, the recognition of cursive scripts using Tesseract is a challenging task. In this paper, Tesseract engine is analyzed and modified for the recognition of Nastalique writing style for Urdu language which is a very complex and cursive writing style of Arabic script. Original Tesseract system has 65.59% and 65.84% accuracies for 14 and 16 font sizes respectively, whereas the modified system, with reduced search space, gives 97.87% and 97.71% accuracies respectively. The efficiency is also improved from an average of 170 milliseconds (ms) to an average of 84 ms for the recognition of Nastalique document images.

Citations

Cited by

Related