vix.ing · top · new · best · stats

A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task Learning

2019/08/15 by Pengfei Wang, Chengquan Zhang, Fei Qi +6 · 82 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Block (permutation group theory) #Boosting (machine learning) #Computer graphics (images) #Computer science #Computer vision #Context (archaeology) #Detector #Graphics #Handwritten Text Recognition Techniques #Image Processing and 3D Reconstruction #Image segmentation #Object detection #Pattern recognition (psychology) #Pixel #Representation (politics) #Segmentation #Single shot #Task (project management) #cs.CV

paper · pdf · doi:10.1145/3343031.3350988

published as In Proceedings of the 27th ACM International Conference on Multimedia (MM '19), October 21-25, 2019, Nice, France · 9 pages, 6 figures, 7 tables, To appear in ACM Multimedia 2019

arxiv created 2019/08/15 · arxiv updated 2019/08/16 · openalex publication_date 2019/10/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Detecting scene text of arbitrary shapes has been a challenging task over the past years. In this paper, we propose a novel segmentation-based text detector, namely SAST, which employs a context attended multi-task learning framework based on a Fully Convolutional Network (FCN) to learn various geometric properties for the reconstruction of polygonal representation of text regions. Taking sequential characteristics of text into consideration, a Context Attention Block is introduced to capture long-range dependencies of pixel information to obtain a more reliable segmentation. In post-processing, a Point-to-Quad assignment method is proposed to cluster pixels into text instances by integrating both high-level object knowledge and low-level pixel information in a single shot. Moreover, the polygonal representation of arbitrarily-shaped text can be extracted with the proposed geometric properties much more effectively. Experiments on several benchmarks, including ICDAR2015, ICDAR2017-MLT, SCUT-CTW1500, and Total-Text, demonstrate that SAST achieves better or comparable performance in terms of accuracy. Furthermore, the proposed algorithm runs at 27.63 FPS on SCUT-CTW1500 with a Hmean of 81.0% on a single NVIDIA Titan Xp graphics card, surpassing most of the existing segmentation-based methods.

Citations