The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
2023/09/21 by Lukas Berglund, Meg Tong, Berglund, Lukas +11 · 11 voices · 140 citations
Computer Science · Psychology · #Business #Curse #Law, AI, and Intellectual Property #Philosophy #Political science #Psychology #Theology #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2309.12288
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/09/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We expose a surprising failure of generalization in auto-regressive large language models (LLMs). If a model is trained on a sentence of the form "A is B", it will not automatically generalize to the reverse direction "B is A". This is the Reversal Curse. For instance, if a model is trained on "Valentina Tereshkova was the first woman to travel to space", it will not automatically be able to answer the question, "Who was the first woman to travel to space?". Moreover, the likelihood of the correct answer ("Valentina Tershkova") will not be higher than for a random name. Thus, models do not generalize a prevalent pattern in their training set: if "A is B" occurs, "B is A" is more likely to occur. It is worth noting, however, that if "A is B" appears in-context, models can deduce the reverse relationship. We provide evidence for the Reversal Curse by finetuning GPT-3 and Llama-1 on fictitious statements such as "Uriah Hawthorne is the composer of Abyssal Melodies" and showing that they fail to correctly answer "Who composed Abyssal Melodies?". The Reversal Curse is robust across model sizes and model families and is not alleviated by data augmentation. We also evaluate ChatGPT (GPT-3.5 and GPT-4) on questions about real-world celebrities, such as "Who is Tom Cruise's mother? [A: Mary Lee Pfeiffer]" and the reverse "Who is Mary Lee Pfeiffer's son?". GPT-4 correctly answers questions like the former 79% of the time, compared to 33% for the latter. Code available at: https://github.com/lukasberglund/reversalcurse.
Cited by
- Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
- The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- Improving Latent Generalization Using Test-time Compute
- Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
- Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
- Whither symbols in the era of advanced neural networks?
- Potemkin Understanding in Large Language Models
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
- Large Language Diffusion Models
- Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
- Evolution and The Knightian Blindspot of Machine Learning
- Connectomics Informed by Large Language Models
- DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
- Cognitive Dark Matter: Measuring What AI Misses
- Fine-Tuned In-Context Learners for Efficient Adaptation
- WeMusic-Agent: Efficient Conversational Music Recommendation via Knowledge Internalization and Agentic Boundary Learning
- Dual-objective Language Models: Training Efficiency Without Overfitting
- Decoding Large Language Diffusion Models with Foreseeing Movement
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- Artificial Intelligence / Human Intelligence: Who Controls Whom?
- PretrainZero: Reinforcement Active Pretraining
- CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- Directional Optimization Asymmetry in Transformers: A Synthetic Stress Test
- Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
- Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- Decomposition of Small Transformer Models
- A mathematical theory of balancing relational generalization and memorization
- When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- Reversal Invariance in Autoregressive Language Models
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems
- Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
- Can Language Models Compose Skills In-Context?
- Code-enabled language models can outperform reasoning models on diverse tasks
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
- From Memorization to Generalization: Fine-Tuning Large Language Models for Biomedical Term-to-Identifier Normalization
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- An Alternative Trajectory for Generative AI
- Abductive Preference Learning
- Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
- From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
- Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
- Watermarking Diffusion Language Models
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- d2Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
- Review of Hallucination Understanding in Large Language and Vision Models
- Bilinear relational structure fixes reversal curse and enables consistent model editing
- SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
- Non-Parametric Structural Priors for Geometry Theorem Prediction
- Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
- Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
- Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- Synthetic bootstrapped pretraining
- Improving LLMs' Learning for Coreference Resolution
- Noise or Nuance: An Investigation Into Useful Information and Filtering For LLM Driven AKBC
- Memorization ≠ Understanding: Do Large Language Models Have the Ability of Scenario Cognition?
- An Investigation on Group Query Hallucination Attacks
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving
- Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
- Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
- Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
- Are LLM Belief Updates Consistent with Bayes' Theorem?
- MeMo: Memory as a Model
- Combinatorial Creativity: A New Frontier in Generalization Abilities
- The benefits of query-based KGQA systems for complex and temporal questions in LLM era
- PropMEND: Hypernetworks for Knowledge Propagation in LLMs
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
- Scaling can lead to compositional generalization
- GPTKB v1.5: A Massive Knowledge Base for Exploring Factual LLM Knowledge
- SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
- Energy-Based Transformers are Scalable Learners and Thinkers
- LEDOM: Reverse Language Model
- Prompting as Scientific Inquiry
- Can Gradient Descent Simulate Prompting?
- Multiple Streams of Knowledge Retrieval: Enriching and Recalling in Transformers
- KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
- LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
- Semantic uncertainty in advanced decoding methods for LLM generation
- Just Go Parallel: Improving the Multilingual Capabilities of Large Language Models
- Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
- Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
- Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
- C-PATH: Conversational Patient Assistance and Triage in Healthcare System
- Quantifying Cross-Modality Memorization in Vision-Language Models
- FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
- VLMs Can Aggregate Scattered Training Patches
- DLM-One: Diffusion Language Models for One-Step Sequence Generation
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis
- Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
- How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering
- Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
- When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
- Causal Cartographer: From Mapping to Reasoning Over Counterfactual Worlds
- Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection
- Data augmentation as a framework for modeling hippocampal contributions to generalization
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
- Memorization-Compression Cycles Improve Generalization
- Multimodal Integrated Knowledge Transfer to Large Language Models through Preference Optimization with Biomedical Applications
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- Me, Myself, and π : Evaluating and Explaining LLM Introspection
- When Role-playing, Do Models Believe What They Say?
- Consistency in Language Models: Current Landscape, Challenges, and Future Directions
- IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
- On the generalization of language models from in-context learning and finetuning: a controlled study
- Memorization and Knowledge Injection in Gated LLMs
- Language models struggle with compartmentalization
- Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
- Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints
- Digital Metabolism: Decoupling Logic from Facts via Regenerative Unlearning -- Towards a Pure Neural Logic Core
- Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmetic
- Capturing Symmetry and Antisymmetry in Language Models through Symmetry-Aware Training Objectives
- AGI Is Coming... Right After AI Learns to Play Wordle
- Exploring the Potential for Large Language Models to Demonstrate Rational Probabilistic Beliefs
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
- Position: It's Time to Optimize LLMs for Self-Consistency
- LMs as Task-Specific Knowledge Bases: An Interpretability Analysis
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Open Problems and a Hypothetical Path Forward in LLM Knowledge Paradigms
- Knowledge-Instruct: Effective Continual Pre-training from Limited Data using Instructions
Discussions
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" (2023) [hn, 25 points, 46 comments]
- Thank you. Have you ever read the Reversal Curse paper? I think there is an interesting parallel to the exploration of prompt engineering as incantation in your preprint. Not because of what the resea [bsky, 5 points, 1 comments]
- KI: Es hapert beim Umkehrschluss [lemmy, 3 points, 4 comments]
- IMO, for the most part they have only tricked people into thinking they are performing logic. I think mostly that's because it's hard to comprehend the amount of training data that has been used. [bsky, 3 points, 1 comments]
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" https://arxiv.org/abs/2309.12288 (https://news.ycombinator.com/item?id=48643184) [bsky, 1 points, 0 comments]
- I don't know if you've seen the work below, which does suggest a difference between the training process and humans memorising stuff that is maybe relevant here (but also, I'm sure an LLM has been pro [bsky, 1 points, 1 comments]
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" https:// arxiv.org/abs/2309.12288 # arxiv # llm # llms [mastodon, 0 points, 0 comments]
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" https://arxiv.org/abs/2309.12288 [bsky, 0 points, 0 comments]
- [2309.12288] The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" — Useful reminder that LLM “knowledge” can be direction-dependent: training on “A is B” doesn’t guarantee it can infer [bsky, 0 points, 0 comments]
- “The reversal curse” (CoRR 2023) arxiv.org/abs/2309.12288 [bsky, 0 points, 0 comments]
- 📰 LLMs trained on "A is B" fail to learn "B is A," highlighting a critical limitation in their ability to generalize and understand bidirectional relationships between concepts, according to a study [bsky, 0 points, 0 comments]
Related