Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
2024/04/08 by Ruiqi Zhang, Licong Lin, Zhang, Ruiqi +5 · 167 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.2404.05868
openalex publication_date 2024/04/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) often memorize sensitive, private, or copyrighted data during pre-training. LLM unlearning aims to eliminate the influence of undesirable data from the pre-trained model while preserving the model's utilities on other tasks. Several practical methods have recently been proposed for LLM unlearning, mostly based on gradient ascent (GA) on the loss of undesirable data. However, on certain unlearning tasks, these methods either fail to effectively unlearn the target data or suffer from catastrophic collapse -- a drastic degradation of the model's utilities. In this paper, we propose Negative Preference Optimization (NPO), a simple alignment-inspired method that could efficiently and effectively unlearn a target dataset. We theoretically show that the progression toward catastrophic collapse by minimizing the NPO loss is exponentially slower than GA. Through experiments on synthetic data and the benchmark TOFU dataset, we demonstrate that NPO-based methods achieve a better balance between unlearning the undesirable data and maintaining the model's utilities. We also observe that NPO-based methods generate more sensible outputs than GA-based methods, whose outputs are often gibberish. Remarkably, on TOFU, NPO-based methods are the first to achieve reasonable unlearning results in forgetting 50% (or more) of the training data, whereas existing methods already struggle with forgetting 10% of training data.
Cited by
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- Understanding Machine Unlearning Through the Lens of Mode Connectivity
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
- Breaking the Curse of Repulsion: Remoteness-Aware Control of Negative Off-Policy Updates
- A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning
- Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
- Feature-Selective Representation Misdirection for Machine Unlearning
- Hard Negative Sample-Augmented DPO Post-Training for Small Language Models
- Sparse Concept Anchoring for Interpretable and Controllable Neural Representations
- When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
- Robust MLLM Unlearning via Visual Knowledge Distillation
- LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
- Grokked Models are Better Unlearners
- Real Time Detection and Quantitative Analysis of Spurious Forgetting in Continual Learning
- What Is Preference Optimization Doing, How and Why?
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- Geometric-disentangelment Unlearning
- Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
- Beyond Superficial Forgetting: Thorough Unlearning through Knowledge Density Estimation and Block Re-insertion
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
- The Realignment Problem: When Right becomes Wrong in LLMs
- A Survey on Unlearning in Large Language Models
- On the Impossibility of Retrain Equivalence in Machine Unlearning
- Probing Knowledge Holes in Unlearned LLMs
- OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- LLM Unlearning with LLM Beliefs
- Not Every Time and Frequency Need to Be Forgotten in Diffusion Unlearning
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
- Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
- Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
- Hierarchical Federated Unlearning for Large Language Models
- Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
- Locket: Robust Feature-Locking Technique for Language Models
- LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
- From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
- SIMU: Selective Influence Machine Unlearning
- LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
- Cross-Modal Attention Guided Unlearning in Vision-Language Models
- Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
- Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
- Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
- Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
- Mitigating Biases in Language Models via Bias Unlearning
- Understanding the Dilemma of Unlearning for Large Language Models
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- Dual-Space Smoothness for Robust and Balanced LLM Unlearning
- Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary
- OFMU: Optimization-Driven Framework for Machine Unlearning
- Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
- CLUE: Conflict-guided Localization for LLM Unlearning Framework
- Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
- Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration
- Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- Module-Aware Parameter-Efficient Machine Unlearning on Transformers
- Apollo: A Posteriori Label-Only Membership Inference Attack Towards Machine Unlearning
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
- Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
- Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
- Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
- CRISP: Persistent Concept Unlearning via Sparse Autoencoders
- Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
- Gradient Surgery for Safe LLM Fine-Tuning
- LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
- LLM Unlearning Without an Expert Curated Dataset
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- IMU: Influence-guided Machine Unlearning
- Towards Evaluation for Real-World LLM Unlearning
- Quantum-Inspired Audio Unlearning: Towards Privacy-Preserving Voice Biometrics
- Unlearning of Knowledge Graph Embedding via Preference Optimization
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
- SUA: Stealthy Multimodal Large Language Model Unlearning Attack
- SoK: Machine Unlearning for Large Language Models
- Memorization Sinks: Isolating Memorization during LLM Training
- Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
- Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
- Unlearning the Noisy Correspondence Makes CLIP More Robust
- When Unlearning Fails: Reliable Data Deletion under Post-Training in Agent Networks
- PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
- LLM Unlearning Should Be Form-Independent
- Revisiting the Past: Data Unlearning with Model State History
- BLUR: A Bi-Level Optimization Approach for LLM Unlearning
- Large Language Model Unlearning for Source Code
- Learning-Time Encoding Shapes Unlearning in LLMs
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
- RULE: Reinforcement UnLEarning Achieves Forget-Retain Pareto Optimality
- Align-then-Unlearn: Embedding Alignment for LLM Unlearning
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
- Improving Large Language Model Safety with Contrastive Representation Learning
- GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
- Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
- Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
- Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language Models
- UNO: Unlearning via Orthogonalization in Generative models
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
- Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
- Rethinking Post-Unlearning Behavior of Large Vision-Language Models
- SALAD: Systematic Assessment of Machine Unlearning on LLM-Aided Hardware Design
- Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
- Speech Unlearning
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
- Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
- Model Unlearning via Sparse Autoencoder Subspace Guided Projections
- Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
- Does Machine Unlearning Truly Remove Knowledge?
- Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
- Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
- BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
- Precise In-Parameter Concept Erasure in Large Language Models
- Machine Unlearning under Overparameterization
- Do Large Language Models (Really) Need Statistical Foundations?
- Soft Weighted Machine Unlearning
- Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models
- Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted data
- Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
- Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
- Losing is for Cherishing: Data Valuation Based on Machine Unlearning and Shapley Value
- UniErase: Towards Balanced and Precise Unlearning in Language Models
- DUSK: Do Not Unlearn Shared Knowledge
- Pre-training Limited Memory Language Models with Internal and External Knowledge
- R-TOFU: Unlearning in Large Reasoning Models
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
- SEPS: A Separability Measure for Robust Unlearning in LLMs
- GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
- Self-Destructive Language Model
- Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
- Ready2Unlearn: A Learning-Time Approach for Preparing Models with Future Unlearning Readiness
- Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
- Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models
- Layered Unlearning for Adversarial Relearning
- ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models
- Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
- LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
- Auditing Language Model Unlearning via Information Decomposition
- Mechanism-Guided Selective Unlearning for RLVR-Induced Reasoning
- Co-LMLM: Continuous-Query Limited Memory Language Models
- Per-parameter Task Arithmetic for Unlearning in Large Language Models
- Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
- Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation
- Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness
- A Model Merging Approach for Continual MLLM Unlearning
- Suppression Sticks, Locality Is Fragile: A Closed-Loop Target-and-Control Audit of Task-Vector Negation in VLA Policies
- DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers
- ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
- A mean teacher algorithm for unlearning of language models
- GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
- SHA256 at SemEval-2025 Task 4: Selective Amnesia -- Constrained Unlearning for Large Language Models via Knowledge Isolation
- GROM: Gradient-Free Rapid One-Shot Machine Unlearning
- Offline Learning and Forgetting for Reasoning with Large Language Models
- LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
- A Neuro-inspired Interpretation of Unlearning in Large Language Models through Sample-level Unlearning Difficulty
- Sharpness-Aware Parameter Selection for Machine Unlearning
Related