Towards Reasoning in Large Language Models: A Survey
2022/12/20 by Jie Huang, Huang, Jie, Kevin Chen–Chuan Chang +1 · 89 citations
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2212.10403
openalex publication_date 2022/12/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Reasoning is a fundamental aspect of human intelligence that plays a crucial role in activities such as problem solving, decision making, and critical thinking. In recent years, large language models (LLMs) have made significant progress in natural language processing, and there is observation that these models may exhibit reasoning abilities when they are sufficiently large. However, it is not yet clear to what extent LLMs are capable of reasoning. This paper provides a comprehensive overview of the current state of knowledge on reasoning in LLMs, including techniques for improving and eliciting reasoning in these models, methods and benchmarks for evaluating reasoning abilities, findings and implications of previous research in this field, and suggestions on future directions. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful discussion and future work.
Cited by
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Language-driven Fine-grained Retrieval
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- Rectifying LLM Thought from Lens of Optimization
- Knowledge Graph Augmented Large Language Models for Disease Prediction
- SelfAI: Building a Self-Training AI System with LLM Agents
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- A perceptual bias of AI Logical Argumentation Ability in Writing
- IPR-1: Interactive Physical Reasoner
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- On the Notion that Language Models Reason
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
- From Natural Language to Certified H-infinity Controllers: Integrating LLM Agents with LMI-Based Synthesis
- Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
- Don't Just Search, Understand: Semantic Path Planning Agent for Spherical Tensegrity Robots in Unknown Environments
- Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
- Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- WeaveRec: An LLM-Based Cross-Domain Sequential Recommendation Framework with Model Merging
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- Can Language Models Compose Skills In-Context?
- Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
- Teaching Language Models to Reason with Tools
- Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
- Reducing Belief Deviation in Reinforcement Learning for Active Reasoning
- SASER: Stego attacks on open-source LLMs
- Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic
- Toward Mechanistic Explanation of Deductive Reasoning in Language Models
- DualResearch: Entropy-Gated Dual-Graph Retrieval for Answer Reconstruction
- In-Context Clustering with Large Language Models
- Modeling Student Learning with 3.8 Million Program Traces
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Making Mathematical Reasoning Adaptive
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- MetaLogic: Robustness Evaluation of Text-to-Image Models via Logically Equivalent Prompts
- MuSLR: Multimodal Symbolic Logical Reasoning
- Generative AI and misinformation: a scoping review of the role of generative AI in the generation, detection, mitigation, and impact of misinformation
- MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- Preference-Based Long-Horizon Robotic Stacking with Multimodal Large Language Models
- LLM DNA: Tracing Model Evolution via Functional Representations
- GEAR: A General Evaluation Framework for Abductive Reasoning
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
- Tracing Uncertainty in Language Model "Reasoning"
- PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning
- On Code-Induced Reasoning in LLMs
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model
- Risk Assessment and Security Analysis of Large Language Models
- Toward PDDL Planning Copilot
- Root Cause Analysis of Radiation Oncology Incidents Using Large Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
- Unveiling the Latent Directions of Reflection in Large Language Models
- Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
- Can Structured Templates Facilitate LLMs in Tackling Harder Tasks? : An Exploration of Scaling Laws by Difficulty
- Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
- Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
- Learning Marked Temporal Point Process Explanations based on Counterfactual and Factual Reasoning
- Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
- AutoIAD: Manager-Driven Multi-Agent Collaboration for Automated Industrial Anomaly Detection
- Open Scene Graphs for Open-World Object-Goal Navigation
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
- Language Model Guided Reinforcement Learning in Quantitative Trading
- HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
- Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- Comparison of Large Language Models for Deployment Requirements
- The Blessing and Curse of Dimensionality in Safety Alignment
Related