2020/11/19 by Gaurav Parmar, Parmar, Gaurav, Dacheng Li +5 · 6 citations
Computer Science · Engineering · #Advanced Image Processing Techniques #Artificial intelligence #Autoencoder #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Deep learning #Digital Media Forensic Detection #Discriminative model #Dual (grammatical number) #Engineering #FOS: Computer and information sciences #Fidelity #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #High fidelity #Image (mathematics) #Inference #Interpolation (computer graphics) #Machine learning #Pattern recognition (psychology) #Representation (politics) #Set (abstract data type) #cs.CV
paper · pdf · doi:10.48550/arxiv.2011.10063
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/11/19 · openalex publication_date 2020/11/19 · arxiv updated 2020/11/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present a new generative autoencoder model with dual contradistinctive losses to improve generative autoencoder that performs simultaneous inference (reconstruction) and synthesis (sampling). Our model, named dual contradistinctive generative autoencoder (DC-VAE), integrates an instance-level discriminative loss (maintaining the instance-level fidelity for the reconstruction/synthesis) with a set-level adversarial loss (encouraging the set-level fidelity for there construction/synthesis), both being contradistinctive. Extensive experimental results by DC-VAE across different resolutions including 32x32, 64x64, 128x128, and 512x512 are reported. The two contradistinctive losses in VAE work harmoniously in DC-VAE leading to a significant qualitative and quantitative performance enhancement over the baseline VAEs without architectural changes. State-of-the-art or competitive results among generative autoencoders for image reconstruction, image synthesis, image interpolation, and representation learning are observed. DC-VAE is a general-purpose VAE model, applicable to a wide variety of downstream tasks in computer vision and machine learning.