vix.ing · top · new · best · stats · spec

Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification

2017/04/15 by Xiang Bai, Mingkun Yang, Bai, Xiang +7 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Convolutional neural network #Embedding #FOS: Computer and information sciences #Focus (optics) #Handwritten Text Recognition Techniques #Image (mathematics) #Image Retrieval and Classification Techniques #Image retrieval #Machine learning #Margin (machine learning) #Natural language processing #Pattern recognition (psychology) #Relevance (law) #Representation (politics) #Semantics (computer science) #Visual Word #Word (group theory) #Word embedding #cs.CV

paper · pdf · doi:10.48550/arxiv.1704.04613

openalex publication_date 2017/04/15 · arxiv created 2017/05/30 · arxiv updated 2017/05/31 · openalex created_date 2019/06/27 · openalex updated_date 2026/08/08

Abstract

Text in natural images contains rich semantics that are often highly relevant to objects or scene. In this paper, we focus on the problem of fully exploiting scene text for visual understanding. The main idea is combining word representations and deep visual features into a globally trainable deep convolutional neural network. First, the recognized words are obtained by a scene text reading system. Then, we combine the word embedding of the recognized words and the deep visual features into a single representation, which is optimized by a convolutional neural network for fine-grained image classification. In our framework, the attention mechanism is adopted to reveal the relevance between each recognized word and the given image, which further enhances the recognition performance. We have performed experiments on two datasets: Con-Text dataset and Drink Bottle dataset, that are proposed for fine-grained classification of business places and drink bottles, respectively. The experimental results consistently demonstrate that the proposed method combining textual and visual cues significantly outperforms classification with only visual representations. Moreover, we have shown that the learned representation improves the retrieval performance on the drink bottle images by a large margin, making it potentially useful in product search.

Citations

Cited by

Related