Toward Mechanistic Explanation of Deductive Reasoning in Language Models
2025/10/10 by Maltoni, Davide, Ferrara, Matteo
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.09340
Abstract
Recent large language models have demonstrated relevant capabilities in solving problems that require logical reasoning; however, the corresponding internal mechanisms remain largely unexplored. In this paper, we show that a small language model can solve a deductive reasoning task by learning the underlying rules (rather than operating as a statistical learner). A low-level explanation of its internal representations and computational circuits is then provided. Our findings reveal that induction heads play a central role in the implementation of the rule completion and rule chaining steps involved in the logical inference required by the task.
Citations
- How do Transformers Learn Implicit Reasoning?
- How Do LLMs Perform Two-Hop Reasoning in Context?
- Logical Reasoning in Large Language Models: A Survey
- JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
- A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning
- MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
- Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
- A Primer on the Inner Workings of Transformer-based Language Models
- How to use and interpret activation patching
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
- A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task
- Transformers, parallel computation, and logarithmic depth
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
- Progress measures for grokking via mechanistic interpretability
- Towards Reasoning in Large Language Models: A Survey
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Probing Classifiers: Promises, Shortcomings, and Advances
- Probing Classifiers: Promises, Shortcomings, and Advances
- Curriculum Learning: A Survey
- Curriculum Learning: A Survey
- ProofWriter: Generating Implications, Proofs, and Abductive Statements\n over Natural Language
- Transformers as Soft Reasoners over Language
Related