vix.ing · top · new · best · stats · spec

Synthesizing Novel Pairs of Image and Text

2017/12/18 by Jason Xie, Xie, Jason, Tingwen Bao +1
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Video Analysis and Summarization #cs.CL #cs.CV #cs.LG

paper · pdf · doi:10.48550/arxiv.1712.06682

arxiv created 2017/12/18 · openalex publication_date 2017/12/18 · arxiv updated 2017/12/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Generating novel pairs of image and text is a problem that combines computer vision and natural language processing. In this paper, we present strategies for generating novel image and caption pairs based on existing captioning datasets. The model takes advantage of recent advances in generative adversarial networks and sequence-to-sequence modeling. We make generalizations to generate paired samples from multiple domains. Furthermore, we study cycles -- generating from image to text then back to image and vise versa, as well as its connection with autoencoders.

Citations

Related