Towards Reasoning in Large Language Models: A Survey
2022/12/20 by Jie Huang, Huang, Jie, Kevin Chen–Chuan Chang +2 · 1 voice · 168 citations
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2212.10403
openalex publication_date 2022/12/20 · arxiv published 2022/12/20 · arxiv updated 2023/05/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Reasoning is a fundamental aspect of human intelligence that plays a crucial role in activities such as problem solving, decision making, and critical thinking. In recent years, large language models (LLMs) have made significant progress in natural language processing, and there is observation that these models may exhibit reasoning abilities when they are sufficiently large. However, it is not yet clear to what extent LLMs are capable of reasoning. This paper provides a comprehensive overview of the current state of knowledge on reasoning in LLMs, including techniques for improving and eliciting reasoning in these models, methods and benchmarks for evaluating reasoning abilities, findings and implications of previous research in this field, and suggestions on future directions. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful discussion and future work.
Cited by
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Language-driven Fine-grained Retrieval
- AGI Requires a Coordination Layer on Top of Pattern Repositories
- Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- CoRT: Code-integrated Reasoning within Thinking
- Rectifying LLM Thought from Lens of Optimization
- Knowledge Graph Augmented Large Language Models for Disease Prediction
- SelfAI: A self-directed framework for long-horizon scientific discovery
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- A perceptual bias of AI Logical Argumentation Ability in Writing
- IPR-1: Interactive Physical Reasoner
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- On the Notion that Language Models Reason
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
- From Natural Language to Certified H-infinity Controllers: Integrating LLM Agents with LMI-Based Synthesis
- Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
- Don't Just Search, Understand: Semantic Path Planning Agent for Spherical Tensegrity Robots in Unknown Environments
- Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
- Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- WeaveRec: An LLM-Based Cross-Domain Sequential Recommendation Framework with Model Merging
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- Can Language Models Compose Skills In-Context?
- Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
- Teaching Language Models to Reason with Tools
- Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- Learning Efficient and Generalizable Graph Retriever for Knowledge-Graph Question Answering
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
- Reducing Belief Deviation in Reinforcement Learning for Active Reasoning
- SASER: Stego attacks on open-source LLMs
- Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic
- Toward Mechanistic Explanation of Deductive Reasoning in Language Models
- DualResearch: Entropy-Gated Dual-Graph Retrieval for Answer Reconstruction
- In-Context Clustering with Large Language Models
- Modeling Student Learning with 3.8 Million Program Traces
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Making Mathematical Reasoning Adaptive
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- MetaLogic: Robustness Evaluation of Text-to-Image Models via Logically Equivalent Prompts
- MuSLR: Multimodal Symbolic Logical Reasoning
- Generative AI and misinformation: a scoping review of the role of generative AI in the generation, detection, mitigation, and impact of misinformation
- MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- Preference-Based Long-Horizon Robotic Stacking with Multimodal Large Language Models
- LLM DNA: Tracing Model Evolution via Functional Representations
- GEAR: A General Evaluation Framework for Abductive Reasoning
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
- Tracing Uncertainty in Language Model "Reasoning"
- PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning
- On Code-Induced Reasoning in LLMs
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model
- Risk Assessment and Security Analysis of Large Language Models
- Toward PDDL Planning Copilot
- Root Cause Analysis of Radiation Oncology Incidents Using Large Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
- Unveiling the Latent Directions of Reflection in Large Language Models
- Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
- Can Structured Templates Facilitate LLMs in Tackling Harder Tasks? : An Exploration of Scaling Laws by Difficulty
- Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
- Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
- Learning Marked Temporal Point Process Explanations based on Counterfactual and Factual Reasoning
- Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
- AutoIAD: Manager-Driven Multi-Agent Collaboration for Automated Industrial Anomaly Detection
- Open Scene Graphs for Open-World Object-Goal Navigation
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
- Language Model Guided Reinforcement Learning in Quantitative Trading
- HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
- Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- Comparison of Large Language Models for Deployment Requirements
- The Blessing and Curse of Dimensionality in Safety Alignment
- Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
- Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
- Knowledge Conceptualization Impacts RAG Efficacy
- ChainEdit: Propagating Ripple Effects in LLM Knowledge Editing through Logical Rule-Guided Chains
- The role of large language models in UI/UX design: A systematic literature review
- LogicGuard: Improving Embodied LLM agents through Temporal Logic based Critics
- Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models
- ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
- Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
- Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
- Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
- Response Quality Assessment for Retrieval-Augmented Generation via Conditional Conformal Factuality
- Baba is LLM: Reasoning in a Game with Dynamic Rules
- Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
- From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
- LLM-based Satisfiability Checking of String Requirements by Consistent Data and Checker Generation
- RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
- MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
- Causes in neuron diagrams, and testing causal reasoning in Large Language Models. A glimpse of the future of philosophy?
- Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
- Long-Short Alignment for Effective Long-Context Modeling in LLMs
- Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making
- Cloud Infrastructure Management in the Age of AI Agents
- OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
- Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
- Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders
- Enhancing Decision-Making of Large Language Models via Actor-Critic
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
- A survey of using EHR as real-world evidence for discovering and validating new drug indications
- Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
- Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
- A Large Language Model-Enabled Control Architecture for Dynamic Resource Capability Exploration in Multi-Agent Manufacturing Systems
- LISRec: Modeling User Preferences with Learned Item Shortcuts for Sequential Recommendation
- FinRipple: Aligning Large Language Models with Financial Market for Event Ripple Effect Awareness
- Exploring the Landscape of Text-to-SQL with Large Language Models: Progresses, Challenges and Opportunities
- Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
- From Reasoning to Learning: A Survey on Hypothesis Discovery and Rule Learning with Large Language Models
- Do LLMs Understand Collaborative Signals? Diagnosis and Repair
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
- Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions
- RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint Data
- The Real Barrier to LLM Agent Usability is Agentic ROI
- Reasoning Meets Personalization: Unleashing the Potential of Large Reasoning Model for Personalized Generation
- Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
- Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
- Children's Mental Models of AI Reasoning: Implications for AI Literacy Education
- Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
- Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
- IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
- Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition
- Interpretable Risk Mitigation in LLM Agent Systems
- Large Language Models for Computer-Aided Design: A Survey
- Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
- Parameterized Argumentation-based Reasoning Tasks for Benchmarking Generative Language Models
- TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- Training Large Reasoning Models Efficiently via Progressive Thought Encoding
- BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
- Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI
- Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
- Generative AI in Education: Student Skills and Lecturer Roles
- L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
- Protoreasoning in Tiny Transformers
- ZeroED: Hybrid Zero-shot Error Detection through Large Language Model Reasoning
- Improving RL Exploration for LLM Reasoning through Retrospective Replay
- Everything You Wanted to Know About LLM-based Vulnerability Detection But Were Afraid to Ask
- Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
- Automating quantum feature map design via large language models
- Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
- Review of Case-Based Reasoning for LLM Agents: Theoretical Foundations, Architectural Components, and Cognitive Integration
- Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey
- PathGPT: Reframing Path Recommendation as a Natural Language Generation Task with Retrieval-Augmented Language Models
Discussions
Related