Transformer Feed-Forward Layers Are Key-Value Memories
2020/12/29 by Mor Geva, Roei Schuster, Geva, Mor +5 · 1 voice · 321 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
paper · pdf · doi:10.48550/arxiv.2012.14913
EMNLP 2021
arxiv published 2020/12/29 · arxiv created 2021/09/05 · arxiv updated 2021/09/07
Abstract
Feed-forward layers constitute two-thirds of a transformer model's parameters, yet their role in the network remains under-explored. We show that feed-forward layers in transformer-based language models operate as key-value memories, where each key correlates with textual patterns in the training examples, and each value induces a distribution over the output vocabulary. Our experiments show that the learned patterns are human-interpretable, and that lower layers tend to capture shallow patterns, while upper layers learn more semantic ones. The values complement the keys' input patterns by inducing output distributions that concentrate probability mass on tokens likely to appear immediately after each pattern, particularly in the upper layers. Finally, we demonstrate that the output of a feed-forward layer is a composition of its memories, which is subsequently refined throughout the model's layers via residual connections to produce the final output distribution.
Cited by
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
- Towards a Relevance Posterior in Neural Information Access
- Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
- Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs
- Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
- Counterfactual Basis Extension and Representational Geometry: An MDL-Constrained Model of Conceptual Growth
- Task Schema and Binding: A Double Dissociation Study of In-Context Learning
- Task Matrices: Linear Maps for Cross-Model Finetuning Transfer
- MIDUS: Memory-Infused Depth Up-Scaling
- Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
- Robust MLLM Unlearning via Visual Knowledge Distillation
- Multi-Granular Node Pruning for Causal Circuit Discovery
- Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
- Interpretation as Linear Transformation: A Cognitive-Geometric Model of Belief and Meaning
- Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
- Flash Multi-Head Feed-Forward Network
- Agency at the Interface: Distinguishing Teleological from Structural Self-Organization via Internal Coarse-Graining and Downward Causation
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- EtCon: Edit-then-Consolidate for Reliable Knowledge Editing
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- Representation Interventions Enable Lifelong Unstructured Knowledge Control
- Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Bridging Philosophy and Machine Learning: A Structuralist Framework for Classifying Neural Network Representations
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
- Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
- CascadedViT: Cascaded Chunk-FeedForward and Cascaded Group Attention Vision Transformer
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- Adaptive Focus Memory for Language Models
- Belief Net: A Filter-Based Framework for Learning Hidden Markov Models from Observations
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- Beyond Superficial Forgetting: Thorough Unlearning through Knowledge Density Estimation and Block Re-insertion
- On the Analogy between Human Brain and LLMs: Spotting Key Neurons in Grammar Perception
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Catching Contamination Before Generation: Spectral Kill Switches for Agents
- Understanding Robustness of Model Editing in Code LLMs
- Balancing Knowledge Updates: Toward Unified Modular Editing in LLMs
- Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
- The Structure of Relation Decoding Linear Operators in Large Language Models
- Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
- Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- MemEIC: A Step Toward Continual and Compositional Knowledge Editing
- A Survey on Unlearning in Large Language Models
- Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
- Verifying Large Language Models' Reasoning Paths via Correlation Matrix Rank
- From Uniform to Adaptive: General Skip-Block Mechanisms for Efficient PDE Neural Operators
- Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
- Probing Neural Combinatorial Optimization Models
- Model-Aware Tokenizer Transfer
- Large Language Models as Model Organisms for Human Associative Learning
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Restoring Pruned Large Language Models via Lost Component Compensation
- DePass: Unified Feature Attributing by Simple Decomposed Forward Pass
- Context-aware Fairness Evaluation and Mitigation in LLMs
- Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Layer Specialization Underlying Compositional Reasoning in Transformers
- Atomic Literary Styling: Mechanistic Manipulation of Prose Generation in Neural Language Models
- Soundness-Aware Level: A Microscopic Signature that Predicts LLM Reasoning Potential
- Emergence of Linear Truth Encodings in Language Models
- Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
- Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
- Simple Projection Variants Improve ColBERT Performance
- KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models
- EvoEdit: Evolving Null-space Alignment for Robust and Efficient Knowledge Editing
- Tapered Language Models
- On the Representations of Entities in Auto-regressive Large Language Models
- Utilizing dynamic sparsity on pretrained DETR
- Understanding the Effects of Domain Finetuning on LLMs
- How to Teach Large Multimodal Models New Skills
- Neuron-Level Analysis of Cultural Understanding in Large Language Models
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
- ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
- Transmuting prompts into weights
- POME: Post Optimization Model Edit via Muon-style Projection
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- MixReasoning: Switching Modes to Think
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- Prototype-Based Dynamic Steering for Large Language Models
- The Geometry of Truth: Layer-wise Semantic Dynamics for Hallucination Detection in Large Language Models
- Mechanistic Interpretability of Socio-Political Frames in Language Models
- What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
- Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
- On Predictability of Reinforcement Learning Dynamics for Large Language Models
- Spectral Scaling Laws in Language Models: How Effectively Do Feed-Forward Networks Use Their Latent Space?
- The Transformer Cookbook
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Sparse Autoencoders Make Audio Foundation Models more Explainable
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Training Dynamics of Parametric and In-Context Knowledge Utilization in Language Models
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Beyond Benchmarks: Understanding Mixture-of-Experts Models through Internal Mechanisms
- Knowledge Homophily in Large Language Models
- Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2
- Detecting (Un)answerability in Large Language Models with Linear Directions
- Multilingual Vision-Language Models, A Survey
- Review of Hallucination Understanding in Large Language and Vision Models
- Fine-tuning Done Right in Model Editing
- Bilinear relational structure fixes reversal curse and enables consistent model editing
- Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- Towards Atoms of Large Language Models
- Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models
- Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
- Diagnosing Model Editing via Knowledge Spectrum
- Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
- Training-free Truthfulness Detection via Sparse MLP Value Vectors
- How Persuasive is Your Context?
- nDNA -- the Semantic Helix of Artificial Cognition
- DISCO: Disentangled Communication Steering for Large Language Models
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- From Parameters to Performance: A Data-Driven Study on LLM Structure and Development
- Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflicts
- Preserving Domain Generalization in Fine-Tuning via Joint Parameter Selection
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms
- Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
- L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
- A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
- Can VLMs Recall Factual Associations From Visual References?
- From Injection to Defense: Constructing Edit-Based Fingerprints for Large Language Models
- Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
- From Confidence to Collapse in LLM Factual Robustness
- Superposition in Graph Neural Networks
- Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision
- No Clustering, No Routing: How Transformers Actually Process Rare Tokens
- When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment
- Universal Properties of Activation Sparsity in Modern Large Language Models
- Mechanistic interpretability for steering vision-language-action models
- LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
- NEAT: Concept driven Neuron Attribution in LLMs
- Scaling Laws for Task-Stratified Knowledge in Post-Training Quantized Large Language Models
- What do language models model? Transformers, automata, and the format of thought
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
- EMSEdit: Efficient Multi-Step Meta-Learning-based Model Editing
- User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
- RegMean++: Enhancing Effectiveness and Generalization of Regression Mean for Model Merging
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- MLP Memory: A Retriever-Pretrained Memory for Large Language Models
- Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations
- RainbowPrompt: Diversity-Enhanced Prompt-Evolving for Continual Learning
- How Chain-of-Thought Works? Tracing Information Flow from Decoding, Projection, and Activation
- CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
- NeuralDB: Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database
- From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- Linear Relational Decoding of Morphology in Language Models
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
- Are Knowledge and Reference in Multilingual Language Models Cross-Lingually Consistent?
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
- Retention analysis of edited knowledge after fine-tuning
- Memorization Sinks: Isolating Memorization during LLM Training
- Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
- Energy-Guided Decoding for Object Hallucination Mitigation
- PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis
- SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs
- Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
- A Survey on Latent Reasoning
- ETT: Expanding the Long Context Understanding Capability of LLMs at Test-Time
- Steering Information Utility in Key-Value Memory for Language Model Post-Training
- LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
- Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
- Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
- Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models
- On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
- Residual Matrix Transformers: Scaling the Size of the Residual Stream
- LLM Unlearning Should Be Form-Independent
- Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
- Multiple Streams of Knowledge Retrieval: Enriching and Recalling in Transformers
- Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models
- Echo State Transformer: Attention Over Finite Memories
- Spark Transformer: Reactivating Sparsity in FFN and Attention
- Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
- Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models
- MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs
- From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
- Can structural correspondences ground real world representational content in Large Language Models?
- PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding
- MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
- UltraSketchLLM: Saliency-Driven Sketching for Ultra-Low Bit LLM Compression
- DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
- PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
- Position: Pause Recycling LoRAs and Prioritize Mechanisms to Uncover Limits and Effectiveness
- Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments
- Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization
- Mitigating Negative Interference in Multilingual Sequential Knowledge Editing through Null-Space Constraints
- Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
- Don't Pay Attention
- Learning Distribution-Wise Control in Representation Space for Language Models
- United Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory
- Defending against Indirect Prompt Injection by Instruction Detection
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
- Attention-Only Transformers via Unrolled Subspace Denoising
- SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
- Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
- Comba: Improving Bilinear RNNs with Closed-loop Control
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
- Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
- How Programming Concepts and Neurons Are Shared in Code Language Models
- Spectral Insights into Data-Oblivious Critical Layers in Large Language Models
- Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
- Exploring the Impact of Occupational Personas on Domain-Specific QA
- Mamba Knockout for Unraveling Factual Information Flow
- Drop Dropout on Single-Epoch Language Model Pretraining
- Disentangling Language and Culture for Evaluating Multilingual Large Language Models
- Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
- In Dialogue with Intelligence: Rethinking Large Language Models as Collective Knowledge
- LoKI: Low-damage Knowledge Implanting of Large Language Models
- SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting
- Precise In-Parameter Concept Erasure in Large Language Models
- Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
- Uncertainty Unveiled: Can Exposure to More In-context Examples Mitigate Uncertainty for Large Language Models?
- Who Reasons in the Large Language Models?
- Pretrained LLMs Learn Multiple Types of Uncertainty
- DocMEdit: Towards Document-Level Model Editing
- SAEs Are Good for Steering -- If You Select the Right Features
- A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models
- REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing
- Benchmarking and Rethinking Knowledge Editing for Large Language Models
- Safety Alignment via Constrained Knowledge Unlearning
- RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection
- Understanding Gated Neurons in Transformers from Their Input-Output Functionality
- Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
- The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
- Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN
- Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
- Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models
- Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
- Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
- How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
- Feed-Forward Steering in Transformer Residual Dynamics
- LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editing
- Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
- Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks
- Void in Language Models
- ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation
- Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
- Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- Truth Neurons
- K-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks
- EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
- A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use
- Optimized Couplings for Watermarking Large Language Models
- On the Superimposed Noise Accumulation Problem in Sequential Knowledge Editing of Large Language Models
- Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition
- UMoE: Unifying Attention and FFN with Shared Experts
- GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models
- Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
- FloE: On-the-Fly MoE Inference on Memory-constrained GPU
- Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
- Scalable LLM Math Reasoning Acceleration with Low-rank Distillation
- Demystifying optimized prompts in language models
- Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models
- LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
- The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
- ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
- The Neuroscience of Transformers
- Epiphany-Aware KV Cache Eviction Without the Attention Matrix
- Gated MLPs as Symmetry-Broken Rank-1 Bilinear Attention
- Variable-Width Transformers
- Adaptive Loops and Memory in Transformers: Think Harder or Know More?
- Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
- On the generalization of language models from in-context learning and finetuning: a controlled study
- STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
- In-Place Test-Time Training
- Polysemy of Synthetic Neurons Towards a New Type of Explanatory Categorical Vector Spaces
- Birdie: Natural Language-Driven Table Discovery Using Differentiable Search Index
- CachePrune: Neural-Based Attribution Defense Against Indirect Prompt Injection Attacks
- Interpretability Can Be Actionable
- Enhancing Vulnerability Reports with Automated and Augmented Description Summarization
- Disentangling MLP Neuron Weights in Vocabulary Space
- Inside the LLM Word Factory
- Neuron Populations Exhibit Divergent Selectivity with Scale
- Factual recall in linear associative memories: sharp asymptotics and mechanistic insights
- Reversible Lifelong Model Editing via Semantic Routing-Based LoRA
- Procedural Pretraining: Warming Up Language Models with Abstract Data
- STEM: Scaling Transformers with Embedding Modules
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
- Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
- Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
- Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model
- DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models
- Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
- Contextual Agentic Memory is a Memo, Not True Memory
- Steering off Course: Reliability Challenges in Steering Language Models
- Exploring How LLMs Capture and Represent Domain-Specific Knowledge
- Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer Enhancement
- Bigram Subnetworks: Mapping to Next Tokens in Transformer Language Models
- Understanding the Repeat Curse in Large Language Models from a Feature Perspective
- CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
- GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
- GROM: Gradient-Free Rapid One-Shot Machine Unlearning
- Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
- Can We Edit LLMs for Long-Tail Biomedical Knowledge?
- How new data permeates LLM knowledge and how to dilute it
- SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling
Discussions
Related