Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
2025/05/19 by Hai Lu, Yueling Liu, Lu, Haolang +11 · 5 citations
Computer Science · #68T27 #Advanced Graph Neural Networks #Adversarial Robustness in Machine Learning #Computers and Society (cs.CY) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.7
paper · pdf · doi:10.48550/arxiv.2505.13143
openalex publication_date 2025/05/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The development of Reasoning Large Language Models (RLLMs) has significantly improved multi-step reasoning capabilities, but it has also made hallucination problems more frequent and harder to eliminate. While existing approaches mitigate hallucinations through external knowledge integration, model parameter analysis, or self-verification, they often fail to capture how hallucinations emerge and evolve across the reasoning chain. In this work, we study the causality of hallucinations under constrained knowledge domains by auditing the Chain-of-Thought (CoT) trajectory and assessing the model's cognitive confidence in potentially erroneous or biased claims. Our analysis reveals that in long-CoT settings, RLLMs can iteratively reinforce biases and errors through flawed reflective reasoning, eventually leading to hallucinated reasoning paths. Surprisingly, even direct interventions at the origin of hallucinations often fail to reverse their effects, as reasoning chains exhibit 'chain disloyalty' -- a resistance to correction and a tendency to preserve flawed logic. Furthermore, we show that existing hallucination detection methods are less reliable and interpretable than previously assumed in complex reasoning scenarios. Unlike methods such as circuit tracing that require access to model internals, our black-box auditing approach supports interpretable long-chain hallucination attribution, offering better generalizability and practical utility. Our code is available at: https://github.com/Winnie-Lian/AHaMetaCognitive
Citations
- HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
- Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
- Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
- On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
- Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Hallucination Mitigation using Agentic AI Natural Language-Based Frameworks
- A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks
- Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
- Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
- Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information
- Evaluation of OpenAI o1: Opportunities and Challenges of AGI
- HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection
- From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
- GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
- Transcoders Find Interpretable LLM Feature Circuits
- Automatically Identifying Local and Global Circuits with Linear Computation Graphs
- Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
- PoLLMgraph: Unraveling Hallucinations in Large Language Models via State Transition Dynamics
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
- Do Large Language Models Latently Perform Multi-Hop Reasoning?
- Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning
- Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- The Impact of Reasoning Step Length on Large Language Models
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- A Survey of Reasoning with Foundation Models
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- FLEEK: Factual Error Detection and Correction with Evidence Retrieved from External Knowledge
- Reflection-Tuning: Data Recycling Improves LLM Instruction-Tuning
- Representation Engineering: A Top-Down Approach to AI Transparency
- Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
- Studying and improving reasoning in humans and machines
- Studying and improving reasoning in humans and machines
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- 🧜Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- How Language Model Hallucinations Can Snowball
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Finding Neurons in a Haystack: Case Studies with Sparse Probing
- Natural Language Reasoning, A Survey
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
- GPT-4 Technical Report
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Language Models (Mostly) Know What They Know
- Survey of Hallucination in Natural Language Generation
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- A Survey on Automated Fact-Checking
- Probing Classifiers: Promises, Shortcomings, and Advances
- Multifaceted Feature Visualization: Uncovering the Different Types of Features Learned By Each Neuron in Deep Neural Networks
Cited by
Related