vix.ing · top · new · best · stats · spec

CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

2022/04/28 by Ming Ding, Ding, Ming, Wendi Zheng +5 · 15 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Video Analysis and Summarization

paper · pdf · doi:10.48550/arxiv.2204.14217

openalex publication_date 2022/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The development of the transformer-based text-to-image models are impeded by its slow generation and complexity for high-resolution images. In this work, we put forward a solution based on hierarchical transformers and local parallel auto-regressive generation. We pretrain a 6B-parameter transformer with a simple and flexible self-supervised task, Cross-modal general language model (CogLM), and finetune it for fast super-resolution. The new text-to-image system, CogView2, shows very competitive generation compared to concurrent state-of-the-art DALL-E-2, and naturally supports interactive text-guided editing on images.

Cited by

Related