A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
2024/01/02 by S. M Towhidul Islam Tonmoy, Tonmoy, S. M Towhidul Islam, S M Mehedi Zaman +11 · 4 voices · 96 citations
Computer Science · Medicine · #Big Data and Digital Economy #COVID-19 diagnosis using AI #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #cs.CL
paper · pdf · doi:10.48550/arxiv.2401.01313
openalex publication_date 2024/01/02 · arxiv published 2024/01/02 · openalex created_date 2024/01/05 · arxiv updated 2024/01/08 · openalex updated_date 2026/07/28
Abstract
As Large Language Models (LLMs) continue to advance in their ability to write human-like text, a key challenge remains around their tendency to hallucinate generating content that appears factual but is ungrounded. This issue of hallucination is arguably the biggest hindrance to safely deploying these powerful LLMs into real-world production systems that impact people's lives. The journey toward widespread adoption of LLMs in practical settings heavily relies on addressing and mitigating hallucinations. Unlike traditional AI systems focused on limited tasks, LLMs have been exposed to vast amounts of online text data during training. While this allows them to display impressive language fluency, it also means they are capable of extrapolating information from the biases in training data, misinterpreting ambiguous prompts, or modifying the information to align superficially with the input. This becomes hugely alarming when we rely on language generation capabilities for sensitive applications, such as summarizing medical records, financial analysis reports, etc. This paper presents a comprehensive survey of over 32 techniques developed to mitigate hallucination in LLMs. Notable among these are Retrieval Augmented Generation (Lewis et al, 2021), Knowledge Retrieval (Varshney et al,2023), CoNLI (Lei et al, 2023), and CoVe (Dhuliawala et al, 2023). Furthermore, we introduce a detailed taxonomy categorizing these methods based on various parameters, such as dataset utilization, common tasks, feedback mechanisms, and retriever types. This classification helps distinguish the diverse approaches specifically designed to tackle hallucination issues in LLMs. Additionally, we analyze the challenges and limitations inherent in these techniques, providing a solid foundation for future research in addressing hallucinations and related phenomena within the realm of LLMs.
Cited by
- Hallucination Detection and Evaluation of Large Language Model
- Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation
- Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting
- Integrating Large Language Models and Knowledge Graphs to Capture Political Viewpoints in News Media
- Calibrated Trust in Dealing with LLM Hallucinations: A Qualitative Study
- PoultryTalk: A Multi-modal Retrieval-Augmented Generation (RAG) System for Intelligent Poultry Management and Decision Support
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- RAG System for Supporting Japanese Litigation Procedures: Faithful Response Generation Complying with Legal Norms
- LLM-Generated Counterfactual Stress Scenarios for Portfolio Risk Simulation via Hybrid Prompt-RAG Pipeline
- The Oracle and The Prism: A Decoupled and Efficient Framework for Generative Recommendation Explanation
- Build AI Assistants using Large Language Models and Agents to Enhance the Engineering Education of Biomechanics
- When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
- Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling
- Evidence-Bound Autonomous Research (EviBound): A Governance Framework for Eliminating False Claims
- PGDA-KGQA: A Prompt-Guided Generative Framework with Multiple Data Augmentation Strategies for Knowledge Graph Question Answering
- Neural Diversity Regularizes Hallucinations in Language Models
- Beyond "Hallucinations": A Framework for Stable Human-AI Reasoning
- LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
- Revisiting Hallucination Detection with Effective Rank-based Uncertainty
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- Large Language Models Hallucination: A Comprehensive Survey
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
- DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability
- CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
- What Should I Cite? A RAG Benchmark for Academic Citation Prediction
- Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
- CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
- Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- Accelerating Discovery: Rapid Literature Screening with LLMs
- Automatic Generation of a Cryptography Misuse Taxonomy Using Large Language Models
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial
- Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
- MeVe: A Modular System for Memory Verification and Effective Context Control in Language Models
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- Toward an Interaction-Centered Approach to Robot Trustworthiness
- An LLM + ASP Workflow for Joint Entity-Relation Extraction
- Incident Response Planning Using a Lightweight Large Language Model with Reduced Hallucination
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
- IM-Chat: A Multi-agent LLM Framework Integrating Tool-Calling and Diffusion Modeling for Knowledge Transfer in Injection Molding Industry
- Artificial intelligence in food safety and nutrition practices: opportunities and risks
- RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
- A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis
- Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- Research on Graph-Retrieval Augmented Generation Based on Historical Text Knowledge Graphs
- Issue Retrieval and Verification Enhanced Supplementary Code Comment Generation
- ConfRAG: Confidence-Guided Retrieval-Augmenting Generation
- Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
- On the Fundamental Impossibility of Hallucination Control in Large Language Models
- Safer Prompts: Reducing Risks from Memorization in Visual Generative AI
- Towards Secure MLOps: Surveying Attacks, Mitigation Strategies, and Research Challenges
- Voice CMS: updating the knowledge base of a digital assistant through conversation
- AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
- ResSVD: Residual Compensated SVD for Large Language Model Compression
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks
- UniErase: Towards Balanced and Precise Unlearning in Language Models
- GAP: Graph-Assisted Prompts for Dialogue-based Medication Recommendation
- Incentivizing Truthful Language Models via Peer Elicitation Games
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
- Truth Neurons
- Eliminating Hallucination-Induced Errors in LLM Code Generation with Functional Clustering
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- Efficient and Reproducible Biomedical Question Answering using Retrieval Augmented Generation
- ”My AI is Lying to Me”: User-reported LLM hallucinations in AI mobile apps reviews
- Chain of Risks Evaluation (CORE): A framework for safer large language models in public mental health
- Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems
- Annotation Vocabulary (Might Be) All You Need
- IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
- Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity
- Grid-Mind: An LLM-Orchestrated Multi-Fidelity Agent for Automated Connection Impact Assessment
- Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
- From Concept to Practice: an Automated LLM-aided UVM Machine for RTL Verification
- Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
- Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
- Guillotine: Hypervisors for Isolating Malicious AIs
- Reflexive Prompt Engineering: A Framework for Responsible Prompt Engineering and Interaction Design
- Sparks of Science: Hypothesis Generation Using Structured Paper Data
- QLLM: Do We Really Need a Mixing Network for Credit Assignment in Multi-Agent Reinforcement Learning?
- A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
- CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
- HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
- Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey
- Generative AI Enhanced Financial Risk Management Information Retrieval
Discussions
Related