vix.ing · top · new · best · stats

The Curious Case of End Token: A Zero-Shot Disentangled Image Editing using CLIP

2024/06/01 by Hidir Yesiltepe, Yesiltepe, Hidir, Yusuf Dalva +3 · 2 citations
Chemistry · Computer Science · #Advanced Neural Network Applications #Artificial intelligence #Chemistry #Computer graphics (images) #Computer science #Computer security #Digital Media Forensic Detection #Generative Adversarial Networks and Image Synthesis #Image (mathematics) #Linguistics #Philosophy #Security token #Shot (pellet) #Zero (linguistics)

paper · pdf · doi:10.48550/arxiv.2406.00457

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Diffusion models have become prominent in creating high-quality images. However, unlike GAN models celebrated for their ability to edit images in a disentangled manner, diffusion-based text-to-image models struggle to achieve the same level of precise attribute manipulation without compromising image coherence. In this paper, CLIP which is often used in popular text-to-image diffusion models such as Stable Diffusion is capable of performing disentangled editing in a zero-shot manner. Through both qualitative and quantitative comparisons with state-of-the-art editing methods, we show that our approach yields competitive results. This insight may open opportunities for applying this method to various tasks, including image and video editing, providing a lightweight and efficient approach for disentangled editing.

Cited by

Related