2021/06/29 by Aviv Gabbay, Gabbay, Aviv, Niv Cohen +3 · 5 citations
Computer Science · #Advanced Image Processing Techniques #Adversarial Robustness in Machine Learning #Algorithm #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Digital Media Forensic Detection #Embedding #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #Identifiability #Image (mathematics) #Machine Learning (cs.LG) #Machine learning #Pattern recognition (psychology) #Residual #Set (abstract data type) #Variation (astronomy) #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2106.15610
published in arXiv (Cornell University) 34 (Cornell University) · NeurIPS 2021. Project page: http://www.vision.huji.ac.il/zerodim
openalex publication_date 2021/06/29 · arxiv created 2021/10/25 · arxiv updated 2021/10/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/08
Unsupervised disentanglement has been shown to be theoretically impossible without inductive biases on the models and the data. As an alternative approach, recent methods rely on limited supervision to disentangle the factors of variation and allow their identifiability. While annotating the true generative factors is only required for a limited number of observations, we argue that it is infeasible to enumerate all the factors of variation that describe a real-world image distribution. To this end, we propose a method for disentangling a set of factors which are only partially labeled, as well as separating the complementary set of residual factors that are never explicitly specified. Our success in this challenging setting, demonstrated on synthetic benchmarks, gives rise to leveraging off-the-shelf image descriptors to partially annotate a subset of attributes in real image domains (e.g. of human faces) with minimal manual effort. Specifically, we use a recent language-image embedding model (CLIP) to annotate a set of attributes of interest in a zero-shot manner and demonstrate state-of-the-art disentangled image manipulation results.