Review of Hallucination Understanding in Large Language and Vision Models
2025/09/26 by Zong-Hua Ho, Ho, Zhengyi, Siyuan Liang +3
Neuroscience · #Artificial Intelligence (cs.AI) #Brain Tumor Detection and Classification #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2510.00034
openalex publication_date 2025/09/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The widespread adoption of large language and vision models in real-world applications has made urgent the need to address hallucinations -- instances where models produce incorrect or nonsensical outputs. These errors can propagate misinformation during deployment, leading to both financial and operational harm. Although much research has been devoted to mitigating hallucinations, our understanding of it is still incomplete and fragmented. Without a coherent understanding of hallucinations, proposed solutions risk mitigating surface symptoms rather than underlying causes, limiting their effectiveness and generalizability in deployment. To tackle this gap, we first present a unified, multi-level framework for characterizing both image and text hallucinations across diverse applications, aiming to reduce conceptual fragmentation. We then link these hallucinations to specific mechanisms within a model's lifecycle, using a task-modality interleaved approach to promote a more integrated understanding. Our investigations reveal that hallucinations often stem from predictable patterns in data distributions and inherited biases. By deepening our understanding, this survey provides a foundation for developing more robust and effective solutions to hallucinations in real-world generative AI systems.
Citations
- Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
- Probabilistic Uncertain Reward Model
- Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
- Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
- Re-evaluating Open-ended Evaluation of Large Language Models
- Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
- Noisy Test-Time Adaptation in Vision-Language Models
- Red-Teaming LLM Multi-Agent Systems via Communication Attacks
- Characterizing Photorealism and Artifacts in Diffusion Model-Generated Images
- Universal Abstraction: Harnessing Frontier Models to Structure Real-World Data at Scale
- Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
- OpenAI's Approach to External Red Teaming for AI Models and Systems
- ICLR: In-Context Learning of Representations
- Chained Tuning Leads to Biased Forgetting
- Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
- Pretrained transformer efficiently learns low-dimensional target functions in-context
- Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
- What is Wrong with Perplexity for Long-context Language Modeling?
- A Geometric Framework for Understanding Memorization in Generative Models
- Representative Social Choice: From Learning Theory to AI Alignment
- A distributional simplicity bias in the learning dynamics of transformers
- Towards Combating Frequency Simplicity-biased Learning for Domain Generalization
- KLay: Accelerating Arithmetic Circuits for Neurosymbolic AI
- Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
- Can In-context Learning Really Generalize to Out-of-distribution Tasks?
- Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
- Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
- Co-occurrence is not Factual Association in Language Models
- Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
- Transformers are Minimax Optimal Nonparametric In-Context Learners
- In-Context Learning with Representations: Contextual Generalization of Trained Transformers
- Detecting, Explaining, and Mitigating Memorization in Diffusion Models
- Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
- LLM Circuit Analyses Are Consistent Across Training and Scale
- Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
- Investigating the Pre-Training Dynamics of In-Context Learning: Task Recognition vs. Task Learning
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
- How Do Large Language Models Acquire Factual Knowledge During Pretraining?
- Transcoders Find Interpretable LLM Feature Circuits
- Understanding Hallucinations in Diffusion Models through Mode Interpolation
- Revisiting Catastrophic Forgetting in Large Language Model Tuning
- Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
- From Unstructured Data to In-Context Learning: Exploring What Tasks Can Be Learned and When
- Language Models Need Inductive Biases to Count Inductively
- Benchmarking and Improving Detail Image Caption
- Learning diverse attacks on large language models for robust red-teaming and safety tuning
- Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data
- On the Noise Robustness of In-Context Learning for Text Generation
- Understanding Forgetting in Continual Learning with Linear Regression
- Looks Too Good To Be True: An Information-Theoretic Analysis of Hallucinations in Generative Restoration Models
- Linking In-context Learning in Transformers to Human Episodic Memory
- Instruction Tuning With Loss Over Instructions
- Base of RoPE Bounds Context Length
- A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
- Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
- Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
- Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
- Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
- LLM Evaluators Recognize and Favor Their Own Generations
- No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
- Bias Amplification in Language Model Evolution: An Iterated Learning Perspective
- Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
- Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
- Instruction-tuned Language Models are Better Knowledge Learners
- Towards Theoretical Understandings of Self-Consuming Generative Models
- In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
- Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction
- Humans or LLMs as the Judge? A Study on Judgement Biases
- MaxMin-RLHF: Alignment with Diverse Human Preferences
- InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
- Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
- ODIN: Disentangled Reward Mitigates Hacking in RLHF
- Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
- A Closer Look at the Limitations of Instruction Tuning
- Panacea: Pareto Alignment via Preference Adaptation for LLMs
- User Intent Recognition and Satisfaction with Large Language Models: A User Study with ChatGPT
- Understanding User Experience in Large Language Model Interactions
- Rethinking FID: Towards a Better Evaluation Metric for Image Generation
- Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop
- LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
- Proving Test Set Contamination in Black Box Language Models
- On the Foundations of Shortcut Learning
- Fidelity-Enriched Contrastive Search: Reconciling the Faithfulness-Diversity Trade-Off in Text Generation
- Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
- Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
- On the Stability of Iterative Retraining of Generative Models on their own Data
- Understanding In-Context Learning from Repetitions
- Benchmarking Cognitive Biases in Large Language Models as Evaluators
- Efficient Streaming Language Models with Attention Sinks
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Elucidating the Exposure Bias in Diffusion Models
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- CausalLM is not optimal for in-context learning
- Kronfluence: Influence Functions with Eigenvalue-corrected Kronecker-Factored Approximate Curvature
- Lost in the Middle: How Language Models Use Long Contexts
- Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks
- Self-Consuming Generative Models Go MAD
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Towards Understanding the Interplay of Generative Artificial Intelligence and the Internet
- Benchmarking Foundation Models with Language-Model-as-an-Examiner
- Exposing Attention Glitches with Flip-Flop Language Modeling
- Understanding and Mitigating Copying in Diffusion Models
- The Impact of Positional Encoding on Length Generalization in Transformers
- Large Language Models are not Fair Evaluators
- Faith and Fate: Limits of Transformers on Compositionality
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Improving Convergence and Generalization Using Parameter Symmetries
- How Spurious Features Are Memorized: Precise Analysis for Random and NTK Features
- Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning
- LIMA: Less Is More for Alignment
- Dissecting Recall of Factual Associations in Auto-Regressive Language Models
- Emergent and Predictable Memorization in Large Language Models
- A Comprehensive Survey on Test-Time Adaptation Under Distribution Shifts
- Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Extracting Training Data from Diffusion Models
- Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models
- Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean Functions
- Large Language Models Struggle to Learn Long-Tail Knowledge
- When and why vision-language models behave like bags-of-words, and what to do about it?
- Why Exposure Bias Matters: An Imitation Learning Perspective of Error Accumulation in Language Generation
- Do Long-Range Language Models Actually Use Long-Range Context?
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Transformer Feed-Forward Layers Are Key-Value Memories
- The Pitfalls of Simplicity Bias in Neural Networks
- Shortcut learning in deep neural networks
- Machine Teaching: A New Paradigm for Building Machine Learning Systems
- Iterative Machine Teaching
- Testing the Manifold Hypothesis
- The serial position effect of free recall.
- Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
- Are Emergent Abilities of Large Language Models a Mirage?
Related