HACK: Hallucinations Along Certainty and Knowledge Axes
2025/10/28 by Simhi, Adi, Herzig, Jonathan, Itzhak, Itay +7 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7
paper · doi:10.48550/arxiv.2510.24222
Abstract
Hallucinations in LLMs present a critical barrier to their reliable usage. Existing research usually categorizes hallucination by their external properties rather than by the LLMs' underlying internal properties. This external focus overlooks that hallucinations may require tailored mitigation strategies based on their underlying mechanism. We propose a framework for categorizing hallucinations along two axes: knowledge and certainty. Since parametric knowledge and certainty may vary across models, our categorization method involves a model-specific dataset construction process that differentiates between those types of hallucinations. Along the knowledge axis, we distinguish between hallucinations caused by a lack of knowledge and those occurring despite the model having the knowledge of the correct response. To validate our framework along the knowledge axis, we apply steering mitigation, which relies on the existence of parametric knowledge to manipulate model activations. This addresses the lack of existing methods to validate knowledge categorization by showing a significant difference between the two hallucination types. We further analyze the distinct knowledge and hallucination patterns between models, showing that different hallucinations do occur despite shared parametric knowledge. Turning to the certainty axis, we identify a particularly concerning subset of hallucinations where models hallucinate with certainty despite having the correct knowledge internally. We introduce a new evaluation metric to measure the effectiveness of mitigation methods on this subset, revealing that while some methods perform well on average, they fail disproportionately on these critical cases. Our findings highlight the importance of considering both knowledge and certainty in hallucination analysis and call for targeted mitigation approaches that consider the hallucination underlying factors.
Citations
- Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations
- AI and the End of an Era
- Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
- Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- The Factuality of Large Language Models in the Legal Domain
- LLMs Will Always Hallucinate, and We Need to Live With This
- WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
- MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
- The Llama 3 Herd of Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Truth is Universal: Robust Detection of Lies in LLMs
- LLM Internal States Reveal Hallucination Risk Faced With a Query
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- To Believe or Not to Believe Your LLM
- Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
- Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question Answering
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- From Persona to Personalization: A Survey on Role-Playing Language Agents
- Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
- Mitigating LLM Hallucinations via Conformal Abstention
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- On Large Language Models' Hallucination with Regard to Known Facts
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
- Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
- Fine-grained Hallucination Detection and Editing for Language Models
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
- Do Androids Know They're Only Dreaming of Electric Sheep?
- The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
- Deficiency of Large Language Models in Finance: An Empirical Examination of Hallucination
- Insights into Classifying and Mitigating LLMs' Hallucinations
- Personas as a Way to Model Truthfulness in Language Models
- AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
- Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- The Troubling Emergence of Hallucination in Large Language Models -- An Extensive Definition, Quantification, and Prescriptive Remediations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
- LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
- Exploring the Relationship between LLM Hallucinations and Prompt Linguistic Nuances: Readability, Formality, and Concreteness
- Quantifying and Attributing the Hallucination of Large Language Models via Association Analysis
- Uncertainty in Natural Language Generation: From Theory to Applications
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation
- Jailbroken: How Does LLM Safety Training Fail?
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Uncertainty in Natural Language Processing: Sources, Quantification, and Applications
- How Language Model Hallucinations Can Snowball
- BloombergGPT: A Large Language Model for Finance
- GPT-4 Technical Report
- Large Language Models Encode Clinical Knowledge
- Large language models encode clinical knowledge
- Prompting GPT-3 To Be Reliable
- Language Models (Mostly) Know What They Know
- The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
- Truthful AI: Developing and governing AI that does not lie
- Measuring and Improving Consistency in Pretrained Language Models
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
- How Can We Know What Language Models Know?
- Language Models as Knowledge Bases?
- Quantifying Uncertainties in Natural Language Processing Tasks
- FEVER: a large-scale dataset for Fact Extraction and VERification
- On Calibration of Modern Neural Networks
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for\n Reading Comprehension
- KBM: Delineating Knowledge Boundary for Adaptive Retrieval in Large Language Models
Cited by
Related