vix.ing · top · new · best · stats · spec

Transformer based Urdu Handwritten Text Optical Character Reader

2022/06/09 by Mohammad Daniyal Shaiq, Shaiq, Mohammad Daniyal, Musa Dildar Ahmed Cheema +3
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Image Processing and 3D Reconstruction #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Vehicle License Plate Recognition

paper · pdf · doi:10.48550/arxiv.2206.04575

openalex publication_date 2022/06/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Extracting Handwritten text is one of the most important components of digitizing information and making it available for large scale setting. Handwriting Optical Character Reader (OCR) is a research problem in computer vision and natural language processing computing, and a lot of work has been done for English, but unfortunately, very little work has been done for low resourced languages such as Urdu. Urdu language script is very difficult because of its cursive nature and change of shape of characters based on it's relative position, therefore, a need arises to propose a model which can understand complex features and generalize it for every kind of handwriting style. In this work, we propose a transformer based Urdu Handwritten text extraction model. As transformers have been very successful in Natural Language Understanding task, we explore them further to understand complex Urdu Handwriting.

Related