vix.ing · top · new · best · stats · spec

DT2I: Dense Text-to-Image Generation from Region Descriptions

2022/04/05 by Stanislav Frolov, Frolov, Stanislav, Prateek Bansal +5
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.2204.02035

openalex publication_date 2022/04/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Despite astonishing progress, generating realistic images of complex scenes remains a challenging problem. Recently, layout-to-image synthesis approaches have attracted much interest by conditioning the generator on a list of bounding boxes and corresponding class labels. However, previous approaches are very restrictive because the set of labels is fixed a priori. Meanwhile, text-to-image synthesis methods have substantially improved and provide a flexible way for conditional image generation. In this work, we introduce dense text-to-image (DT2I) synthesis as a new task to pave the way toward more intuitive image generation. Furthermore, we propose DTC-GAN, a novel method to generate images from semantically rich region descriptions, and a multi-modal region feature matching loss to encourage semantic image-text matching. Our results demonstrate the capability of our approach to generate plausible images of complex scenes using region captions.

Related