Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
2026/07/08 by Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.CR #cs.LG #cs.MM
paper · pdf · doi:10.18653/v1/2026.findings-acl.1379
Abstract
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data. Retraining after deletion requests or policy updates is often impractical, and targeted forgetting remains difficult because knowledge is distributed across shared representations. Multimodal unlearning addresses this challenge by enabling selective removal across modalities while retaining overall utility. This survey offers a unified, system-oriented view of multimodal unlearning across vision, language, audio, and video, grounded in recent advances, emerging applications, and open problems. Our taxonomy enables systematic comparison across model architectures and modalities, clarifying trade-offs among deletion strength, retention, efficiency, reversibility, and robustness. This survey highlights open problems and practical considerations to support future research and deployment of multimodal unlearning. We release a curated repository: https://smsnobin77.github.io/Awesome-Multimodal-Unlearning/
Citations
- Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
- No Encore: Unlearning as Opt-Out in Music Generation
- UnGuide: Learning to Forget with LoRA-Guided Diffusion Models
- LoReUn: Data Itself Implicitly Provides Cues to Improve Machine Unlearning
- Quantum-Inspired Audio Unlearning: Towards Privacy-Preserving Voice Biometrics
- Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
- Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
- Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning
- Concept Unlearning by Modeling Key Steps of Diffusion Process
- Unlearning the Noisy Correspondence Makes CLIP More Robust
- PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
- Membership Inference Attacks as Privacy Tools: Reliability, Disparity and Ensemble
- SUA: Stealthy Multimodal Large Language Model Unlearning Attack
- Video Unlearning via Low-Rank Refusal Vector
- "Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
- CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models
- Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
- The Dual Power of Interpretable Token Embeddings: Jailbreaking Attacks and Defenses for Diffusion Model Unlearning
- Backdoor Defense in Diffusion Models via Spatial Attention Unlearning
- Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware Optimization
- Prompting Forgetting: Unlearning in GANs via Textual Guidance
- Human Motion Unlearning
- Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimization
- Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
- A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
- SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning
- SEMU: Singular Value Decomposition for Efficient Machine Unlearning
- SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
- Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
- Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters
- Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
- VLSBench: Unveiling Visual Leakage in Multimodal Safety
- Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset
- Model Integrity when Unlearning with T2I Diffusion Models
- CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
- Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
- CLEAR: Character Unlearning in Textual and Visual Modalities
- Dynamic Negative Guidance of Diffusion Models
- SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
- UnSeg: One Universal Unlearnable Example Generator is Enough against All Image Segmentation
- Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
- Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning
- Efficient Backdoor Defense in Multimodal Contrastive Learning: A Token-Level Unlearning Method for Mitigating Threats
- Score Forgetting Distillation: A Swift, Data-Free Method for Machine Unlearning in Diffusion Models
- Unlearning or Concealment? A Critical Analysis and Evaluation Metrics for Unlearning in Diffusion Models
- Enhancing User-Centric Privacy Protection: An Interactive Framework through Diffusion Models and Machine Unlearning
- DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
- Controllable Unlearning for Image-to-Image Generative Models via ε-Constrained Optimization
- Machine Unlearning in Generative AI: A Survey
- Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning
- Targeted Unlearning with Single Layer Unlearning Gradient
- On Large Language Model Continual Unlearning
- Zero-Shot Class Unlearning in CLIP with Synthetic Samples
- DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
- MU-Bench: A Multitask Multimodal Benchmark for Machine Unlearning
- SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
- Large Language Model Unlearning via Embedding-Corrupted Prompts
- MUC: Machine Unlearning for Contrastive Learning with Black-box Evaluation
- Multi-Modal Recommendation Unlearning for Legal, Licensing, and Modality Constraints
- Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient
- FreezeAsGuard: Mitigating Illegal Adaptation of Diffusion Models via Selective Tensor Freezing
- Probing Unlearned Diffusion Models: A Transferable Adversarial Attack Perspective
- Evaluating Text-to-Visual Generation with Image-to-Text Generation
- CPR: Retrieval Augmented Generation for Copyright Protection
- Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning
- Hiding and Recovering Knowledge in Text-to-Image Diffusion Models via Learnable Prompts
- MACE: Mass Concept Erasure in Diffusion Models
- Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning
- Machine Unlearning for Image-to-Image Generative Models
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
- Unified Concept Editing in Diffusion Models
- EmoSet: A Large-scale Visual Emotion Dataset with Rich Attributes
- Training Data Attribution for Diffusion Models
- SneakyPrompt: Jailbreaking Text-to-image Generative Models
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
- Model Sparsity Can Simplify Machine Unlearning
- CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive Learning
- Unlearnable Clusters: Towards Label-agnostic Unlearnable Examples
- Pre-trained Encoders in Self-Supervised Learning Improve Secure and Privacy-preserving Supervised Learning
- Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models
- DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
- DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
- Measuring the Carbon Intensity of AI in Cloud Instances
- Membership Inference Attacks From First Principles
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Remember What You Want to Forget: Algorithms for Machine Unlearning
- Diffusion Earth Mover's Distance and Distribution Embeddings
- Descent-to-Delete: Gradient-Based Methods for Machine Unlearning
- Machine Unlearning
- Certified Data Removal from Machine Learning Models
- Making AI Forget You: Data Deletion in Machine Learning
- OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
- A Style-Based Generator Architecture for Generative Adversarial Networks
- A Style-Based Generator Architecture for Generative Adversarial Networks
- A Corpus for Reasoning About Natural Language Grounded in Photographs
- Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias
- AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark
- Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
- VizWiz Grand Challenge: Answering Visual Questions from Blind People
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Demystifying MMD GANs
- Progressive Growing of GANs for Improved Quality, Stability, and Variation
- VGGFace2: A dataset for recognising faces across pose and age
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- Membership Inference Attacks against Machine Learning Models
- Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature
- VQA: Visual Question Answering
- Learning Face Representation from Scratch
- Deep Learning Face Attributes in the Wild
- Microsoft COCO: Common Objects in Context
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Calibrating Noise to Sensitivity in Private Data Analysis
- Hierarchy-Aware Multimodal Unlearning for Medical AI
- UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models
- Bridging Language and Items for Retrieval and Recommendation
- Gradient-based learning applied to document recognition
Discussions
Related