2025/09/25 by Nicola Novello, Novello, Nicola, Fontana, Federico +6
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #Control Systems and Identification #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG) #Matrix Theory and Algorithms
paper · pdf · doi:10.48550/arxiv.2509.21167
openalex publication_date 2025/09/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Most existing methods for concept unlearning in text-to-image diffusion models minimize a mean squared error (MSE) loss between the denoiser outputs conditioned on a target and an anchor concept, which is implicitly the KL divergence between two Gaussians. We generalize this objective to any f-divergence, recovering MSE as the KL instance, and identify a family of α-divergences whose Gaussian closed-form yields cheap, MSE-like training objectives. For the remaining f-divergences, we provide a min-max objective based on the variational formulation of the f-divergence. We theoretically analyze and numerically validate how different f-divergences impact the gradient magnitude and the convergence properties of the algorithm, affecting the quality of unlearning. For instance, we observe that the Hellinger closed-form instance consistently dominates MSE across multiple scenarios. More generally, the proposed unified framework offers a flexible paradigm for selecting the optimal divergence based on the application and user goal, allowing for finer control over the trade-off between unlearning efficacy and generative fidelity.