vix.ing · top · new · best · stats

Image and Encoded Text Fusion for Multi-Modal Classification

2018/10/03 by Ignazio Gallo, Gallo, Ignazio, Alessandro Calefati +5 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications #cs.CV

paper · pdf · doi:10.48550/arxiv.1810.02001

Accepted to DICTA 2018

arxiv created 2018/10/03 · openalex publication_date 2018/10/03 · arxiv updated 2018/10/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Multi-modal approaches employ data from multiple input streams such as textual and visual domains. Deep neural networks have been successfully employed for these approaches. In this paper, we present a novel multi-modal approach that fuses images and text descriptions to improve multi-modal classification performance in real-world scenarios. The proposed approach embeds an encoded text onto an image to obtain an information-enriched image. To learn feature representations of resulting images, standard Convolutional Neural Networks (CNNs) are employed for the classification task. We demonstrate how a CNN based pipeline can be used to learn representations of the novel fusion approach. We compare our approach with individual sources on two large-scale multi-modal classification datasets while obtaining encouraging results. Furthermore, we evaluate our approach against two famous multi-modal strategies namely early fusion and late fusion.

Cited by

Related