Metacognition in LLMs: Foundations, Progress, and Opportunities
2026/07/13 by Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu +3 · 1 voice
Computer Science · #cs.CL #cs.AI
paper · pdf
Abstract
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs' metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion. An organized list of papers can be found at https://github.com/yale-nlp/LLM-Metacognition.
Citations
- Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
- LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
- Reasoning Models Generate Societies of Thought
- The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance
- Emergent Introspective Awareness in Large Language Models
- Do Large Language Models Know What They Are Capable Of?
- MDToC: Metacognitive Dynamic Tree of Concepts for Boosting Mathematical Problem-Solving of Large Language Models
- Evaluating Large Language Models in Scientific Discovery
- Metacognitive Sensitivity for Test-Time Dynamic Model Selection
- Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
- Monitor-Generate-Verify (MGV): Formalising Metacognitive Theory for Language Model Reasoning
- Metacognition and Confidence Dynamics in Advice Taking from Generative AI
- Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
- Before you , monitor: Implementing Flavell's metacognitive framework in LLMs
- Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- Improving Metacognition and Uncertainty Communication in Language Models
- What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
- Evidence for Limited Metacognition in LLMs
- Agentic Metacognition: Designing a "Self-Aware" Low-Code Agent for Failure Prediction and Human Handoff
- Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
- Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
- Language Models Coupled with Metacognition Can Outperform Reasoning Models
- Meta-R1: Empowering Large Reasoning Models with Metacognition
- ObjexMT: Objective Extraction and Metacognitive Calibration for LLM-as-a-Judge under Multi-Turn Jailbreaks
- Privileged Self-Access Matters for Introspection in AI
- Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents
- MeLA: A Metacognitive LLM-Driven Architecture for Automatic Heuristic Design
- Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
- Quantifying uncert-AI-nty: Testing the accuracy of LLMs’ confidence judgments
- CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs
- MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Towards Understanding the Cognitive Habits of Large Reasoning Models
- Query-Level Uncertainty in Large Language Models
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
- Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
- Does It Make Sense to Speak of Introspection in Large Language Models?
- Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
- Are Reasoning Models More Prone to Hallucination?
- Enhancing Critical Thinking in Generative AI Search with Metacognitive Prompts
- Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition
- Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
- When Two LLMs Debate, Both Think They'll Win
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
- ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Theory of Mind in Large Language Models: Assessment and Enhancement
- AI Awareness
- Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
- Metacognition and Uncertainty Communication in Humans and Large Language Models
- DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
- Understanding R1-Zero-Like Training: A Critical Perspective
- ThinkPatterns-21k: A Systematic Study on the Impact of Thinking Patterns in LLMs
- Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
- ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
- Language Models Fail to Introspect About Their Knowledge of Language
- Learning from Failures in Multi-Attempt Reinforcement Learning
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction
- Re-evaluating Theory of Mind evaluation in large language models
- A Survey of Uncertainty Estimation Methods on Large Language Models
- Do Large Language Models Know How Much They Know?
- Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger
- Which Type of Students can LLMs Act? Investigating Authentic Simulation with Graph-based Human-AI Collaborative System
- Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
- FLOCAL: Architecture for Metacognitive Autotelic Artificial Intelligence
- State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence
- Tell me about yourself: LLMs are aware of their learned behaviors
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
- LegalAgentBench: Evaluating LLM Agents in Legal Domain
- Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
- On Verbalized Confidence Scores for LLMs
- Alignment faking in large language models
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance
- Pragmatic Metacognitive Prompting Improves LLM Performance on Sarcasm Detection
- Competence-Aware AI Agents with Metacognition for Unknown Situations and Environments (MUSE)
- A Survey on Human-Centric LLMs
- Imagining and building wise machines: The centrality of AI metacognition
- Addressing Uncertainty in LLMs to Enhance Reliability in Generative AI
- Improving Uncertainty Quantification in Large Language Models via Semantic Embeddings
- A Survey of Uncertainty Estimation in LLMs: Theory Meets Practice
- Judgment of Learning: A Human Ability Beyond Generative Artificial Intelligence
- Looking Inward: Language Models Can Learn About Themselves by Introspection
- Taming Overconfidence in LLMs: Reward Calibration in RLHF
- Performance and Metacognition Disconnect when Reasoning in Human-AI Interaction
- Large Language Models for Disease Diagnosis: A Scoping Review
- Metacognitive Myopia in Large Language Models
- Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
- Metacognitive AI: Framework and the Case for a Neurosymbolic Approach
- Meta Reasoning for Large Language Models
- Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models
- Cycles of Thought: Measuring LLM Confidence through Stable Explanations
- To Believe or Not to Believe Your LLM
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
- Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
- Devil's Advocate: Anticipatory Reflection for LLM Agents
- Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving
- Generative AI as a metacognitive agent: A comparative mixed-method study with human participants on ICF-mimicking exam performance
- Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
- Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
- Metacognitive Retrieval-Augmented Large Language Models
- What large language models know and what people think they know
- Metacognition is all you need? Using Introspection in Generative Agents to Improve Goal-directed Behavior
- Escalation Risks from Language Models in Military and Diplomatic Decision-Making
- MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
- The Metacognitive Demands and Opportunities of Generative AI
- Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling
- Violation of Expectation via Metacognitive Prompting Reduces Theory of Mind Prediction Error in Large Language Models
- Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness
- Metacognitive Prompting Improves Understanding in Large Language Models
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Do Large Language Models Know What They Don't Know?
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- MoT: Memory-of-Thought Enables ChatGPT to Self-Improve
- Large Linguistic Models: Investigating LLMs' metalinguistic abilities
- Poisoning Language Models During Instruction Tuning
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
- Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models
- Poisoning Web-Scale Training Datasets is Practical
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Language Models (Mostly) Know What They Know
- On the probability-quality paradox in language generation
- Truthful AI: Developing and governing AI that does not lie
- From internal models toward metacognitive AI
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- On Calibration of Modern Neural Networks
- Metacognition and consciousness
- Unskilled and unaware of it: How difficulties in recognizing one's own incompetence lead to inflated self-assessments.
- Consciousness and metacognition.
- MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
- Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
- From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control
- Enhancing Mathematical Reasoning in the Classroom: The Effects of Cooperative Learning and Metacognitive Training
- A metacognitive view of individual differences in self-regulated learning
- Metacognition and cognitive monitoring: A new area of cognitive–developmental inquiry.
- VERIFICATION OF FORECASTS EXPRESSED IN TERMS OF PROBABILITY
Discussions