Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
2024/04/08 by Ruiqi Zhang, Licong Lin, Zhang, Ruiqi +5 · 81 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.2404.05868
openalex publication_date 2024/04/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) often memorize sensitive, private, or copyrighted data during pre-training. LLM unlearning aims to eliminate the influence of undesirable data from the pre-trained model while preserving the model's utilities on other tasks. Several practical methods have recently been proposed for LLM unlearning, mostly based on gradient ascent (GA) on the loss of undesirable data. However, on certain unlearning tasks, these methods either fail to effectively unlearn the target data or suffer from catastrophic collapse -- a drastic degradation of the model's utilities. In this paper, we propose Negative Preference Optimization (NPO), a simple alignment-inspired method that could efficiently and effectively unlearn a target dataset. We theoretically show that the progression toward catastrophic collapse by minimizing the NPO loss is exponentially slower than GA. Through experiments on synthetic data and the benchmark TOFU dataset, we demonstrate that NPO-based methods achieve a better balance between unlearning the undesirable data and maintaining the model's utilities. We also observe that NPO-based methods generate more sensible outputs than GA-based methods, whose outputs are often gibberish. Remarkably, on TOFU, NPO-based methods are the first to achieve reasonable unlearning results in forgetting 50% (or more) of the training data, whereas existing methods already struggle with forgetting 10% of training data.
Cited by
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- Understanding Machine Unlearning Through the Lens of Mode Connectivity
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
- Breaking the Curse of Repulsion: Remoteness-Aware Control of Negative Off-Policy Updates
- A Mechanistic Perspective and Difficulty Metric for Unlearning
- Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
- Feature-Selective Representation Misdirection for Machine Unlearning
- Hard Negative Sample-Augmented DPO Post-Training for Small Language Models
- Sparse Concept Anchoring for Interpretable and Controllable Neural Representations
- When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
- Robust MLLM Unlearning via Visual Knowledge Distillation
- LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
- Grokked Models are Better Unlearners
- Real Time Detection and Quantitative Analysis of Spurious Forgetting in Continual Learning
- What Is Preference Optimization Doing, How and Why?
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- Geometric-disentangelment Unlearning
- Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
- Beyond Superficial Forgetting: Thorough Unlearning through Knowledge Density Estimation and Block Re-insertion
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
- The Realignment Problem: When Right becomes Wrong in LLMs
- A Survey on Unlearning in Large Language Models
- On the Impossibility of Retrain Equivalence in Machine Unlearning
- Probing Knowledge Holes in Unlearned LLMs
- OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- LLM Unlearning with LLM Beliefs
- Not Every Time and Frequency Need to Be Forgotten in Diffusion Unlearning
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
- Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
- Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
- Hierarchical Federated Unlearning for Large Language Models
- Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
- Locket: Robust Feature-Locking Technique for Language Models
- LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
- From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
- SIMU: Selective Influence Machine Unlearning
- LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
- Cross-Modal Attention Guided Unlearning in Vision-Language Models
- Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
- Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
- Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
- Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
- Mitigating Biases in Language Models via Bias Unlearning
- Understanding the Dilemma of Unlearning for Large Language Models
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- Dual-Space Smoothness for Robust and Balanced LLM Unlearning
- Decision Potential Surface: A Theoretical and Practical Approximation of LLM's Decision Boundary
- OFMU: Optimization-Driven Framework for Machine Unlearning
- Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
- CLUE: Conflict-guided Localization for LLM Unlearning Framework
- Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
- Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration
- Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- Module-Aware Parameter-Efficient Machine Unlearning on Transformers
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
- Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
- Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
- Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
- CRISP: Persistent Concept Unlearning via Sparse Autoencoders
- Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
- Gradient Surgery for Safe LLM Fine-Tuning
- LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
- LLM Unlearning Without an Expert Curated Dataset
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- IMU: Influence-guided Machine Unlearning
- Towards Evaluation for Real-World LLM Unlearning
- Quantum-Inspired Audio Unlearning: Towards Privacy-Preserving Voice Biometrics
- Unlearning of Knowledge Graph Embedding via Preference Optimization
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
Related