vix.ing · top · new · best · stats · spec

Zero-Shot Text-to-Image Generation

2021/02/24 by Aditya Ramesh, Ramesh, Aditya, Mikhail Pavlov +13 · 2 voices · 346 citations
Computer Science · #Advanced Neural Network Applications #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #cs.CV #cs.LG

paper · pdf · doi:10.48550/arxiv.2102.12092

openalex publication_date 2021/02/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.

Citations

Cited by

Discussions

Related