2021/07/26 by Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Bhunia, Ayan Kumar +5 · 2 citations
Computer Science · Mathematics · #Affine transformation #Artificial intelligence #Autoencoder #Character (mathematics) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Deep learning #Disjoint sets #FOS: Computer and information sciences #Feature (linguistics) #Focus (optics) #Handwritten Text Recognition Techniques #Image Retrieval and Classification Techniques #Inference #Machine learning #Mathematics #Salient #Space (punctuation) #Transformation (genetics) #Video Analysis and Summarization #cs.CV
paper · pdf · doi:10.48550/arxiv.2107.12081
published in arXiv (Cornell University) (Cornell University) · IEEE International Conference on Computer Vision (ICCV), 2021
arxiv created 2021/07/26 · openalex publication_date 2021/07/26 · arxiv updated 2021/07/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Visual text recognition is undoubtedly one of the most extensively researched topics in computer vision. Great progress have been made to date, with the latest models starting to focus on the more practical "in-the-wild" setting. However, a salient problem still hinders practical deployment -- prior arts mostly struggle with recognising unseen (or rarely seen) character sequences. In this paper, we put forward a novel framework to specifically tackle this "unseen" problem. Our framework is iterative in nature, in that it utilises predicted knowledge of character sequences from a previous iteration, to augment the main network in improving the next prediction. Key to our success is a unique cross-modal variational autoencoder to act as a feedback module, which is trained with the presence of textual error distribution data. This module importantly translate a discrete predicted character space, to a continuous affine transformation parameter space used to condition the visual feature map at next iteration. Experiments on common datasets have shown competitive performance over state-of-the-arts under the conventional setting. Most importantly, under the new disjoint setup where train-test labels are mutually exclusive, ours offers the best performance thus showcasing the capability of generalising onto unseen words.