vix.ing · top · new · best · stats

Dual Adversarial Inference for Text-to-Image Synthesis

2019/08/14 by Qicheng Lao, Lao, Qicheng, Mohammad Havaei +9 · 5 citations
Computer Science · #Adversarial system #Artificial intelligence #Composition (language) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Dual (grammatical number) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Image (mathematics) #Image synthesis #Inference #Linguistics #Multimodal Machine Learning Applications #Natural language processing #Process (computing) #Quality (philosophy) #Space (punctuation) #Style (visual arts) #Video Analysis and Summarization #cs.CV

paper · pdf · doi:10.48550/arxiv.1908.05324

published in arXiv (Cornell University) (Cornell University) · Accepted to ICCV 2019

arxiv created 2019/08/14 · openalex publication_date 2019/08/14 · arxiv updated 2019/08/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Synthesizing images from a given text description involves engaging two types of information: the content, which includes information explicitly described in the text (e.g., color, composition, etc.), and the style, which is usually not well described in the text (e.g., location, quantity, size, etc.). However, in previous works, it is typically treated as a process of generating images only from the content, i.e., without considering learning meaningful style representations. In this paper, we aim to learn two variables that are disentangled in the latent space, representing content and style respectively. We achieve this by augmenting current text-to-image synthesis frameworks with a dual adversarial inference mechanism. Through extensive experiments, we show that our model learns, in an unsupervised manner, style representations corresponding to certain meaningful information present in the image that are not well described in the text. The new framework also improves the quality of synthesized images when evaluated on Oxford-102, CUB and COCO datasets.

Citations

Related