2018/12/22 by Zhilin Zheng, Zheng, Zhilin, Li Sun +1 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
paper · pdf · doi:10.48550/arxiv.1812.09502
openalex publication_date 2018/12/22 · openalex created_date 2019/06/27 · openalex updated_date 2026/07/28
VAE requires the standard Gaussian distribution as a prior in the latent space. Since all codes tend to follow the same prior, it often suffers the so-called "posterior collapse". To avoid this, this paper introduces the class specific distribution for the latent code. But different from CVAE, we present a method for disentangling the latent space into the label relevant and irrelevant dimensions, \bmzs and \bmzu, for a single input. We apply two separated encoders to map the input into \bmzs and \bmzu respectively, and then give the concatenated code to the decoder to reconstruct the input. The label irrelevant code \bmzu represent the common characteristics of all inputs, hence they are constrained by the standard Gaussian, and their encoder is trained in amortized variational inference way, like VAE. While \bmzs is assumed to follow the Gaussian mixture distribution in which each component corresponds to a particular class. The parameters for the Gaussian components in \bmzs encoder are optimized by the label supervision in a global stochastic way. In theory, we show that our method is actually equivalent to adding a KL divergence term on the joint distribution of \bmzs and the class label c, and it can directly increase the mutual information between \bmzs and the label c. Our model can also be extended to GAN by adding a discriminator in the pixel domain so that it produces high quality and diverse images.