Faith and Fate: Limits of Transformers on Compositionality
2023/05/29 by Nouha Dziri, Ximing Lu, Dziri, Nouha +29 · 13 voices · 135 citations
Computer Science · Materials Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2305.18654
openalex publication_date 2023/05/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Transformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems. This begs the question: Are these errors incidental, or do they signal more substantial limitations? In an attempt to demystify transformer LLMs, we investigate the limits of these models across three representative compositional tasks -- multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures. Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with increased task complexity.
Cited by
- Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam
- Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation
- Cognitive Dark Matter: Measuring What AI Misses
- Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
- Curriculum Guided Massive Multi Agent System Solving For Robust Long Horizon Tasks
- AGI Requires a Coordination Layer on Top of Pattern Repositories
- AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation
- Nexus: Higher-Order Attention Mechanisms in Transformers
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models
- In-Context Compositional Learning via Sparse Coding Transformer
- Understanding the Staged Dynamics of Transformers in Learning Latent Structure
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- Cognitive Maps in Language Models: A Mechanistic Analysis of Spatial Planning
- Frontier Large Language Models Rival State-of-the-Art Planners
- The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
- Next-Latent Prediction Transformers Learn Compact World Models
- DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
- Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation
- The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
- Out-of-distribution generalization via composition: A lens through induction heads in Transformers
- When No Paths Lead to Rome: Benchmarking Systematic Neural Relational Reasoning
- Can Language Models Compose Skills In-Context?
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- Evaluating LLM Reasoning Beyond Correctness and CoT
- DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- UniCode: Augmenting Evaluation for Code Reasoning
- Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languages
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- An Alternative Trajectory for Generative AI
- How Do Language Models Compose Functions?
- RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
- Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
- Generating Meaning: Active Inference and the Scope and Limits of Passive AI
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
- When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education
- A model of errors in transformers
- Review of Hallucination Understanding in Large Language and Vision Models
- Teaching Transformers to Solve Combinatorial Problems through Efficient Trial & Error
- From Found to Designed: Concepts as a Design Axis for Large Language Models
- SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge
- Generating meaning: active inference and the scope and limits of passive AI
- Generative AI for Economic Research: Use Cases and Implications for Economists
- Non-Parametric Structural Priors for Geometry Theorem Prediction
- Understanding Subword Compositionality of Large Language Models
- Variation in Verification: Understanding Verification Dynamics in Large Language Models
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
- Transforming Agency. On the mode of existence of Large Language Models
- Provable Benefits of In-Tool Learning for Large Language Models
- Dream 7B: Diffusion Large Language Models
- TransLLM: A Unified Multi-Task Foundation Framework for Urban Transportation via Learnable Prompting
- Reinforced Context Order Recovery for Adaptive Reasoning and Planning
- Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- The Missing Reward: Active Inference in the Era of Experience
- Topos Theory for Generative AI and LLMs
- Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time
- CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
- Towards Consistent Long-Term Pose Generation
- Adaptive Multi-Agent Reasoning via Automated Workflow Generation
- Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs
- Scaling can lead to compositional generalization
- Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs
- The Thin Line Between Comprehension and Persuasion in LLMs
- Frontiers of Generative AI for Network Optimization: Theories, Limits, and Visions
- Discourse Heuristics For Paradoxically Moral Self-Correction
- Not All Explanations for Deep Learning Phenomena Are Equally Valuable
- Improving Large Language Models with Concept-Aware Fine-Tuning
- Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
- π-CoT: Prolog-Initialized Chain-of-Thought Prompting for Multi-Hop Question-Answering
- Long-Context Generalization with Sparse Attention
- LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
- Distinct Computations Emerge From Compositional Curricula in In-Context Learning
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
- Learning to Insert [PAUSE] Tokens for Better Reasoning
- Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
- On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures
- Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
- Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
- The Road to Generalizable Neuro-Symbolic Learning Should be Paved with Foundation Models
- Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models
- Scalable Complexity Control Facilitates Reasoning Ability of LLMs
- Learning Composable Chains-of-Thought
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- GIFARC: Synthetic Dataset for Leveraging Human-Intuitive Analogies to Elevate AI Reasoning
- Revisiting Bi-Linear State Transitions in Recurrent Neural Networks
- Two Causally Related Needles in a Video Haystack
- Recursive Decomposition with Dependencies for Generic Divide-and-Conquer Reasoning
- Learning Extrapolative Sequence Transformations from Markov Chains
- A Theoretical Analysis of Compositional Generalization in Neural Networks: A Necessary and Sufficient Condition
- Unraveling Misinformation Propagation in LLM Reasoning
- The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- Small Models, Smarter Learning: The Power of Joint Task Training
- Knot So Simple: A Minimalistic Environment for Spatial Reasoning
- Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
- Why Large Language Models Fail at Tabular Prediction
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
- A Minimum Description Length Approach to Regularization in Neural Networks
- Systematic Generalization in Language Models Scales with Information Entropy
- Language and Thought: The View from LLMs
- Enhancing Latent Computation in Transformers with Latent Tokens
- No Consciousness? No Meaning (and no AGI!)
- PoE-World: Compositional World Modeling with Products of Programmatic Experts
- Lost in Transmission: When and Why LLMs Fail to Reason Globally
- Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions
- Chain-of-Thought Tokens are Computer Program Variables
- When Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?
- The limits of large language models and the necessity of human cognition in K-12 education
- Arithmetic Pedagogy for Language Models
- APWA: A Distributed Architecture for Parallelizable Agentic Workflows
- Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation
- Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
- SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
- Rethinking Reflection in Pre-Training
- Contemplative Agent
- Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
Discussions
- I'd distinguish "spurious correlation" shortcuts from "next-token shortcuts" . The former is independent of next-token training: in an "x+y=z" task depending on how you distribute your x and y, some [bsky, 4 points, 1 comments]
- Faith and Fate: Limits of Transformers on Compositionality [hn, 3 points, 1 comments]
- Here's my quick review and thoughts of Dziri et al.'s recent paper on the limits of auto-regressive language model capabilities. arxiv.org/abs/2305.18654 [bsky, 2 points, 8 comments]
- Faith and Fate - Limits of Transformers on Compositionality [lemmy, 2 points, 0 comments]
- Faith and Fate: Limits of Transformers on Compositionality [hn, 1 points, 0 comments]
- Faith and Fate: Limits of Transformers on Compositionality [hn, 1 points, 0 comments]
- Faith and Fate
arxiv.org/pdf/2305.18654
@nouhadziri.bsky.social
presents experiments suggesting that LLMs do not learn the basic (generalizable) capability to compose partial solutions of composition [bsky, 1 points, 1 comments]
- AIs will inevitably fail and depart from the real plane as they are used for more and more complex tasks. arxiv.org/abs/2305.18654 [bsky, 1 points, 0 comments]
- Faith and Fate: Limits of Transformers on Compositionality [hn, 1 points, 0 comments]
- Perhaps they would consider having a look at this paper? arxiv.org/abs/2305.18654 [bsky, 0 points, 0 comments]
- That paper he cites in that slide is great arxiv.org/abs/2305.18654 [bsky, 0 points, 0 comments]
- Faith and Fate: Limits of Transformers on Compositionality [bsky, 0 points, 1 comments]
- どうもTransformerなどの自己回帰型のLLMでは多段階の抽象的な推論は難しい感じがする。前の算数計算の組み合わせ推論の研究 arxiv.org/abs/2305.18654 を見てもそういう気しかしない… [bsky, 0 points, 0 comments]
Related