Mass-Editing Memory in a Transformer
2022/10/13 by Kevin Meng, Meng, Kevin, Arnab Sen Sharma +7 · 116 citations
Computer Science · Decision Sciences · #Topic Modeling #Advanced Data Storage Technologies #Scientific Computing and Data Management
paper · pdf · doi:10.48550/arxiv.2210.07229
Abstract
Recent work has shown exciting promise in updating large language models with new memories, so as to replace obsolete information or add specialized knowledge. However, this line of work is predominantly limited to updating single associations. We develop MEMIT, a method for directly updating a language model with many memories, demonstrating experimentally that it can scale up to thousands of associations for GPT-J (6B) and GPT-NeoX (20B), exceeding prior work by orders of magnitude. Our code and data are at https://memit.baulab.info.
Cited by
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
- Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
- Forgetting Is Not a Fix: Path Dependence in Sequential Engram Editing
- TokenMem: Faithful Knowledge Injection for Frozen LLMs
- Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs
- Investigating Model Editing for Unlearning in Large Language Models
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
- An Information-Theoretic Framework for Robust Large Language Model Editing
- Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
- SALVE: Sparse Autoencoder-Latent Vector Editing for Mechanistic Control of Neural Networks
- Task Matrices: Linear Maps for Cross-Model Finetuning Transfer
- Towards Effective Model Editing for LLM Personalization
- Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
- LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
- Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
- The Road of Adaptive AI for Precision in Cybersecurity
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- EtCon: Edit-then-Consolidate for Reliable Knowledge Editing
- EvoEdit: Lifelong Free-Text Knowledge Editing through Latent Perturbation Augmentation and Knowledge-driven Parameter Fusion
- RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
- Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge Editing
- Lightweight Model Editing for LLMs to Correct Deprecated API Recommendations
- Representation Interventions Enable Lifelong Unstructured Knowledge Control
- Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
- Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
- BlockCert: Certified Blockwise Extraction of Transformer Mechanisms
- Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
- Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
- Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
- SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised Attention
- Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
- Understanding Robustness of Model Editing in Code LLMs
- Balancing Knowledge Updates: Toward Unified Modular Editing in LLMs
- Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
- Localized Adaptation Reveals Distinct Learning Signatures in Transformers
- ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- Understanding Multi-View Transformers
- EditMark: Watermarking Large Language Models based on Model Editing
- OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
- Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
- Dynamic Retriever for In-Context Knowledge Editing via Policy Optimization
- The Impact of Negated Text on Hallucination with Large Language Models
- KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning
- MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
- DSCD: Large Language Model Detoxification with Self-Constrained Decoding
- CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
- STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models
- EvoEdit: Evolving Null-space Alignment for Robust and Efficient Knowledge Editing
- REPAIR: Robust Editing via Progressive Adaptive Intervention and Reintegration
- On the Representations of Entities in Auto-regressive Large Language Models
- Neuron-Level Analysis of Cultural Understanding in Large Language Models
- ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
- SIMU: Selective Influence Machine Unlearning
- POME: Post Optimization Model Edit via Muon-style Projection
- Energy-Regularized Sequential Model Editing on Hyperspheres
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
- Is Model Editing Built on Sand? Revealing Its Illusory Success and Fragile Foundation
- Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
- Are Robust LLM Fingerprints Adversarially Robust?
- EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
- SUIT: Knowledge Editing with Subspace-Aware Key-Value Mappings
- Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
- Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2
- Fine-tuning Done Right in Model Editing
- Bilinear relational structure fixes reversal curse and enables consistent model editing
- A Unified Framework for Diffusion Model Unlearning with f-Divergence
- Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
- bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI
- MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
- Consistency-Aware Parameter-Preserving Knowledge Editing Framework for Multi-Hop Question Answering
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
- Diagnosing Model Editing via Knowledge Spectrum
- Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
- Digging Into the Internal: Causality-Based Analysis of LLM Function Calling
- Preserving Domain Generalization in Fine-Tuning via Joint Parameter Selection
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms
- Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition
- Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA
- PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement
- Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
- Model Science: getting serious about verification, explanation and control of AI systems
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
- Delta-Audit: Explaining What Changes When Models Change
- Castle: Causal Cascade Updates in Relational Databases with Large Language Models
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Side Effects of Erasing Concepts from Diffusion Models
- Evaluating Sparse Autoencoders for Monosemantic Representation
- CRISP: Persistent Concept Unlearning via Sparse Autoencoders
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- ALAS: Autonomous Learning Agent for Self-Updating Language Models
- A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
- DySK-Attn: A Framework for Efficient, Real-Time Knowledge Updating in Large Language Models via Dynamic Sparse Knowledge Attention
- Surgical Knowledge Rewrite in Compact LLMs: An 'Unlearn-then-Learn' Strategy with (IA3) for Localized Factual Modulation and Catastrophic Forgetting Mitigation
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models
- EMSEdit: Efficient Multi-Step Meta-Learning-based Model Editing
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
- Aligning Language Models with Real-time Knowledge Editing
- Latent Knowledge Scalpel: Precise and Massive Knowledge Editing for Large Language Models
- Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
- Knowledge Editing for Multi-Hop Question Answering Using Semantic Analysis
- Modular Delta Merging with Orthogonal Constraints: A Scalable Framework for Continual and Reversible Model Composition
- MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System Theory
- NeuralDB: Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database
Related