2021/03/31 by Amrutha Saseendran, Saseendran, Amrutha, Kathrin Skubch +3
Computer Science · #Generative Adversarial Networks and Image Synthesis #Digital Media Forensic Detection #Image Processing and 3D Reconstruction
paper · pdf · doi:10.48550/arxiv.2103.16795
Image generation has rapidly evolved in recent years. Modern architectures\nfor adversarial training allow to generate even high resolution images with\nremarkable quality. At the same time, more and more effort is dedicated towards\ncontrolling the content of generated images. In this paper, we take one further\nstep in this direction and propose a conditional generative adversarial network\n(GAN) that generates images with a defined number of objects from given\nclasses. This entails two fundamental abilities (1) being able to generate\nhigh-quality images given a complex constraint and (2) being able to count\nobject instances per class in a given image. Our proposed model modularly\nextends the successful StyleGAN2 architecture with a count-based conditioning\nas well as with a regression sub-network to count the number of generated\nobjects per class during training. In experiments on three different datasets,\nwe show that the proposed model learns to generate images according to the\ngiven multiple-class count condition even in the presence of complex\nbackgrounds. In particular, we propose a new dataset, CityCount, which is\nderived from the Cityscapes street scenes dataset, to evaluate our approach in\na challenging and practically relevant scenario.\n