Cognitive Foundations for Reasoning and Their Manifestation in LLMs
2025/11/20 by Priyanka Kargupta, Shuyue Stella Li, Kargupta, Priyanka +23 · 1 voice · 3 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Multimodal Machine Learning Applications #Topic Modeling #cs.AI
paper · pdf · doi:10.48550/arxiv.2511.16660
openalex publication_date 2025/11/20 · openalex created_date 2025/11/23 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) solve complex problems yet fail on simpler variants, suggesting they achieve correct outputs through mechanisms fundamentally different from human reasoning. To understand this gap, we synthesize cognitive science research into a taxonomy of 28 cognitive elements spanning reasoning invariants, meta-cognitive controls, representations for organizing reasoning & knowledge, and transformation operations. We introduce a fine-grained evaluation framework and conduct the first large-scale empirical analysis of 192K traces from 18 models across text, vision, and audio, complemented by 54 human think-aloud traces, which we make publicly available. We find that models under-utilize cognitive elements correlated with success, narrowing to rigid sequential processing on ill-structured problems where diverse representations and meta-cognitive monitoring are critical. Human traces show more abstraction and conceptual processing, while models default to surface-level enumeration. Meta-analysis of 1.6K LLM reasoning papers reveals the research community concentrates on easily quantifiable elements (sequential organization: 55%, decomposition: 60%) but neglecting meta-cognitive controls (self-awareness: 16%) that correlate with success. Models possess behavioral repertoires associated with success but fail to deploy them spontaneously. Leveraging these patterns, we develop test-time reasoning guidance that automatically scaffold successful structures, improving performance by up to 66.7% on complex problems. By establishing a shared vocabulary between cognitive science and LLM research, our framework enables systematic diagnosis of reasoning failures and principled development of models that reason through robust cognitive mechanisms rather than spurious shortcuts, while providing tools to test theories of human cognition at scale.
Citations
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
- Personalized Reasoning: Just-In-Time Personalization and Why LLMs Fail At It
- Fluid Language Model Benchmarking
- Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
- PrefPalette: Personalized Preference Modeling with Latent Attributes
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study
- Potemkin Understanding in Large Language Models
- Spurious Rewards: Rethinking Training Signals in RLVR
- Beyond True or False: Retrieval-Augmented Hierarchical Analysis of Nuanced Claims
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
- To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
- Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models
- DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
- Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2
- s1: Simple test-time scaling
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
- A Survey on Large Language Models for Code Generation
- From Frege to chatGPT: Compositionality in language, cognition, and deep neural networks
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
- The pitfalls of next-token prediction
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
- Let's Verify Step by Step
- Faith and Fate: Limits of Transformers on Compositionality
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Language Model Behavior: A Comprehensive Survey
- Language Model Behavior: A Comprehensive Survey
- Progress measures for grokking via mechanistic interpretability
- A fine-grained comparison of pragmatic language understanding in humans and language models
- Solving math word problems with process- and outcome-based feedback
- Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs
- Language Models are Multilingual Chain-of-Thought Reasoners
- Bidirectional Language Models Are Also Few-shot Learners
- Emergent Abilities of Large Language Models
- Large Language Models are Zero-Shot Reasoners
- Can language models learn from explanations in context?
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- Probing Classifiers: Promises, Shortcomings, and Advances
- Probing Classifiers: Promises, Shortcomings, and Advances
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- The Language of Generalization
- The Large‐Scale Structure of Semantic Networks: Statistical Analyses and a Model of Semantic Growth
- Constraint relaxation and chunk decomposition in insight problem solving.
- Cognitive Load During Problem Solving: Effects on Learning
- Analogical problem solving
- Family resemblances: Studies in the internal structure of categories
- Vacunación del niño inmigrante y adoptado en España
- Metacognition and cognitive monitoring: A new area of cognitive–developmental inquiry.
- Perception in chess
Cited by
Discussions
Related