2019/12/20 by Manuel Carbonell, Carbonell, Manuel, Alícia Fornés +5
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1912.10016
openalex publication_date 2019/12/20 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
In the last years, the consolidation of deep neural network architectures for\ninformation extraction in document images has brought big improvements in the\nperformance of each of the tasks involved in this process, consisting of text\nlocalization, transcription, and named entity recognition. However, this\nprocess is traditionally performed with separate methods for each task. In this\nwork we propose an end-to-end model that combines a one stage object detection\nnetwork with branches for the recognition of text and named entities\nrespectively in a way that shared features can be learned simultaneously from\nthe training error of each of the tasks. By doing so the model jointly performs\nhandwritten text detection, transcription, and named entity recognition at page\nlevel with a single feed forward step. We exhaustively evaluate our approach on\ndifferent datasets, discussing its advantages and limitations compared to\nsequential approaches. The results show that the model is capable of benefiting\nfrom shared features for simultaneously solving interdependent tasks.\n