Linearity of Relation Decoding in Transformer Language Models
2023/08/17 by Evan Hernandez, Arnab Sen Sharma, Hernandez, Evan +13 · 1 voice · 46 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2308.09124
openalex publication_date 2023/08/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Much of the knowledge encoded in transformer language models (LMs) may be expressed in terms of relations: relations between words and their synonyms, entities and their attributes, etc. We show that, for a subset of relations, this computation is well-approximated by a single linear transformation on the subject representation. Linear relation representations may be obtained by constructing a first-order approximation to the LM from a single prompt, and they exist for a variety of factual, commonsense, and linguistic relations. However, we also identify many cases in which LM predictions capture relational knowledge accurately, but this knowledge is not linearly encoded in their representations. Our results thus reveal a simple, interpretable, but heterogeneously deployed knowledge representation strategy in transformer LMs.
Cited by
- Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks
- Distributed Sparse Interventions in Language Models
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Task Matrices: Linear Maps for Cross-Model Finetuning Transfer
- Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
- The Structure of Relation Decoding Linear Operators in Large Language Models
- Sequences of Logits Reveal the Low Rank Structure of Language Models
- STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models
- On the Entity-Level Alignment in Crosslingual Consistency
- How Do Language Models Compose Functions?
- Training Dynamics of Parametric and In-Context Knowledge Utilization in Language Models
- Bilinear relational structure fixes reversal curse and enables consistent model editing
- STEREODISCO: Discovering Stereotypicality in LLMs
- Do Activation Verbalization Methods Convey Privileged Information?
- Quantifying Compositionality of Classic and State-of-the-Art Embeddings
- All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
- Beyond Transcription: Mechanistic Interpretability in ASR
- Linear Relational Decoding of Morphology in Language Models
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
- Linearly Decoding Refused Knowledge in Aligned Language Models
- Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
- Multiple Streams of Knowledge Retrieval: Enriching and Recalling in Transformers
- Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
- Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
- Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
- Context-Robust Knowledge Editing for Language Models
- On the Emergence of Linear Analogies in Word Embeddings
- Disentangling Knowledge Representations for Large Language Model Editing
- Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
- Through a Compressed Lens: Investigating the Impact of Quantization on LLM Explainability and Interpretability
- Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
- Relation Extraction or Pattern Matching? Unravelling the Generalisation Limits of Language Models for Biographical RE
- Do different prompting methods yield a common task representation in language models?
- Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
- Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
- Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
- LLM Safety From Within: Detecting Harmful Content with Internal Representations
- Differential syntactic and semantic encoding in LLMs
- Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models
- Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
- Exploring How LLMs Capture and Represent Domain-Specific Knowledge
- Functional Abstraction of Knowledge Recall in Large Language Models
- On Linear Representations and Pretraining Data Frequency in Language Models
- LMs as Task-Specific Knowledge Bases: An Interpretability Analysis
Discussions
Related