A Survey of Hallucination in Large Foundation Models
2023/09/12 by Vipula Rawte, Amit Sheth, Rawte, Vipula +3 · 1 voice · 89 citations
Computer Science · Economics, Econometrics and Finance · Neuroscience · Psychology · #Artificial Intelligence (cs.AI) #Complex Systems and Time Series Analysis #Computation and Language (cs.CL) #FOS: Computer and information sciences #Functional Brain Connectivity Studies #Information Retrieval (cs.IR) #Mental Health Research Topics #cs.AI #cs.CL #cs.IR
paper · pdf · doi:10.48550/arxiv.2309.05922
openalex publication_date 2023/09/12 · arxiv published 2023/09/12 · arxiv updated 2023/09/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Hallucination in a foundation model (FM) refers to the generation of content that strays from factual reality or includes fabricated information. This survey paper provides an extensive overview of recent efforts that aim to identify, elucidate, and tackle the problem of hallucination, with a particular focus on ``Large'' Foundation Models (LFMs). The paper classifies various types of hallucination phenomena that are specific to LFMs and establishes evaluation criteria for assessing the extent of hallucination. It also examines existing strategies for mitigating hallucination in LFMs and discusses potential directions for future research in this area. Essentially, the paper offers a comprehensive examination of the challenges and solutions related to hallucination in LFMs.
Cited by
- FaithLens: Detecting and Explaining Faithfulness Hallucination
- Multi-Agent LLM Committees for Autonomous Software Beta Testing
- Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
- A Categorical Analysis of Large Language Models and Why LLMs Circumvent the Symbol Grounding Problem
- Can LLMs Make (Personalized) Access Control Decisions?
- Enhancing Sequential Recommendation with World Knowledge from Large Language Models
- Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
- What Limits Agentic Systems Efficiency?
- Spurious reconstruction from brain activity
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- A Survey on Evaluation of Large Language Models
- Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Domain-Grounded Evaluation of LLMs in International Student Knowledge
- TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching
- Navigating the AI era: university communication strategies and perspectives on generative AI tools
- Advances in Large Language Models for Medicine
- The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
- A funny companion: Distinct neural responses to AI- versus human-attributed humor
- MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- Select to Know: An Internal-External Knowledge Self-Selection Framework for Domain-Specific Question Answering
- Can we Evaluate RAGs with Synthetic Data?
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
- GOFAI meets Generative AI: Development of Expert Systems by means of Large Language Models
- SCALEFeedback: A Large-Scale Dataset of Synthetic Computer Science Assignments for LLM-generated Educational Feedback Research
- Dean of LLM Tutors: A Framework for Automated Quality Review of AI-generated Feedback
- Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
- Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
- Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration
- First Hallucination Tokens Are Different from Conditional Ones
- Hallucination Detection and Mitigation with Diffusion in Multi-Variate Time-Series Foundation Models
- Mitigating Object Hallucinations via Sentence-Level Early Intervention
- Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding
- Multi-Agent Retrieval-Augmented Framework for Evidence-Based Counterspeech Against Health Misinformation
- Technical Report for Argoverse2 Scenario Mining Challenges on Iterative Error Correction and Spatially-Aware Prompting
- Structural Decoupling: A Scaffold-Flow Theory of Generalization and Alignment
- Self-Critique-Guided Curiosity Refinement: Enhancing Honesty and Helpfulness in Large Language Models via In-Context Learning
- HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models
- On hallucinations in AI-generated content for nuclear medicine imaging (the DREAM report)
- A Framework for Generating Conversational Recommendation Datasets from Behavioral Interactions
- Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
- Defending against Indirect Prompt Injection by Instruction Detection
- Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
- Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
- Machine Mirages: Defining the Undefined
- Invariant Link Selector for Spatial-Temporal Out-of-Distribution Problem
- RISE: Reasoning Enhancement via Iterative Self-Exploration in Multi-hop Question Answering
- Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap
- ChartLens: Fine-grained Visual Attribution in Charts
- Writing Like the Best: Exemplar-Based Expository Text Generation
- Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs
- Learning Interpretable Representations Leads to Semantically Faithful EEG-to-Text Generation
- Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs
- Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
- Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
- Concept Incongruence: An Exploration of Time and Death in Role Playing
- The Role of Visualization in LLM-Assisted Knowledge Graph Systems: Effects on User Trust, Exploration, and Workflows
- Policy Contrastive Decoding for Robotic Foundation Models
- Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
- Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild
- Efficient and Scalable Neural Symbolic Search for Knowledge Graph Complex Query Answering
- BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model Editing
- Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity
- Multimodal Large Language Models for Medicine: A Comprehensive Survey
- A Generative-AI-Driven Claim Retrieval System Capable of Detecting and Retrieving Claims from Social Media Platforms in Multiple Languages
- Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction
- Credible Plan-Driven RAG Method for Multi-Hop Question Answering
- Transforming remanufacturing automation with large language models: A forward-looking analysis with case studies
- Towards Visual Text Grounding of Multimodal Large Language Model
- ArXivBench: When You Should Avoid Using ChatGPT for Academic Writing
- Insights from Verification: Training a Verilog Generation LLM with Reinforcement Learning with Testbench Feedback
- FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation
- Pricing AI Model Accuracy
- Rethinking Technological Readiness in the Era of AI Uncertainty
- MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework
- Enhancing Large Language Models through Neuro-Symbolic Integration and Ontological Reasoning
- How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
- Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey
Discussions
Related