A Survey of Multimodal Hallucination Evaluation and Detection
2025/07/25 by Zhiyuan Chen, Yuecong Min, Chen, Zhiyuan +11 · 4 citations
Social Sciences · Computer Science · #Misinformation and Its Impacts #Adversarial Robustness in Machine Learning #Digital Media Forensic Detection
paper · pdf · doi:10.48550/arxiv.2507.19024
Abstract
Multi-modal Large Language Models (MLLMs) have emerged as a powerful paradigm for integrating visual and textual information, supporting a wide range of multi-modal tasks. However, these models often suffer from hallucination, producing content that appears plausible but contradicts the input content or established world knowledge. This survey offers an in-depth review of hallucination evaluation benchmarks and detection methods across Image-to-Text (I2T) and Text-to-image (T2I) generation tasks. Specifically, we first propose a taxonomy of hallucination based on faithfulness and factuality, incorporating the common types of hallucinations observed in practice. Then we provide an overview of existing hallucination evaluation benchmarks for both T2I and I2T tasks, highlighting their construction process, evaluation objectives, and employed metrics. Furthermore, we summarize recent advances in hallucination detection methods, which aims to identify hallucinated content at the instance level and serve as a practical complement of benchmark-based evaluation. Finally, we highlight key limitations in current benchmarks and detection methods, and outline potential directions for future research.
Citations
- A Low-Rank Method for Vision Language Model Hallucination Mitigation in Autonomous Driving
- Human-AI Collaborative Uncertainty Quantification
- Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
- Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
- Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- AgriVLN: Vision-and-Language Navigation for Agricultural Robots
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
- HalLoc: Token-level Localization of Hallucinations for Vision Language Models
- Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
- VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
- DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
- Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
- Multimodal Large Language Models for Medicine: A Comprehensive Survey
- VLLFL: A Vision-Language Model Based Lightweight Federated Learning Framework for Smart Agriculture
- Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks
- TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
- EgoBlind: Towards Egocentric Visual Assistance for the Blind
- Exploring Bias in over 100 Text-to-Image Generative Models
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- Treble Counterfactual VLMs: A Causal Approach to Hallucination
- Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias
- See What You Are Told: Visual Attention Sink in Large Multimodal Models
- Vision-Language Models Struggle to Align Entities across Modalities
- MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
- CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
- Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
- Efficient Diffusion Models: A Survey
- WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
- Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
- Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent
- T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
- DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
- FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models
- Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
- Leapfrog Latent Consistency Model (LLCM) for Medical Images Generation
- VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
- Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
- V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
- A Survey of Hallucination in Large Visual Language Models
- Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
- LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
- Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
- Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
- Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
- Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models
- NeIn: Telling What You Don't Want
- V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models
- Reference-free Hallucination Detection for Large Vision-Language Models
- Leveraging Vision Language Models for Specialized Agricultural Tasks
- Multi-Object Hallucination in Vision-Language Models
- Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
- Evaluating the Quality of Hallucination Benchmarks for Large Vision-Language Models
- VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
- PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
- AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
- Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
- Understanding Hallucinations in Diffusion Models through Mode Interpolation
- Stealthy Targeted Backdoor Attacks against Image Captioning
- Calibrated Self-Rewarding Vision Language Models
- VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes
- THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
- Hallucination of Multimodal Large Language Models: A Survey
- Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
- Efficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models Via Diffusion Models
- Lossy Image Compression with Foundation Diffusion Models
- Tackling Structural Hallucination in Image Translation with Local Diffusion
- No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
- Evaluating Text-to-Visual Generation with Image-to-Text Generation
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Visual Hallucination: Definition, Quantification, and Prescriptive Remediations
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
- PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
- B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
- RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
- HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
- Visual Hallucinations of Multi-modal Large Language Models
- Synthesizing CTA Image Data for Type-B Aortic Dissection using Stable Diffusion Models
- ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
- Unified Hallucination Detection for Multimodal Large Language Models
- A Survey on Hallucination in Large Vision-Language Models
- The Neglected Tails in Vision-Language Models
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- Deep Learning-based Image and Video Inpainting: A Survey
- EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question Answering
- Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites
- Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
- Rethinking FID: Towards a Better Evaluation Metric for Image Generation
- OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
- Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
- Mitigating Hallucination in Visual Language Models with Visual Supervision
- HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
- SpectralGPT: Spectral Remote Sensing Foundation Model
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Don't Make Your LLM an Evaluation Benchmark Cheater
- FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip Connection
- ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
- Creative Robot Tool Use with Large Language Models
- Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
- Improved Baselines with Visual Instruction Tuning
- HallE-Control: Controlling Object Hallucination in Large Multimodal Models
- Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
- Aligning Large Multimodal Models with Factually Augmented RLHF
- MUTEX: Learning Unified Policies from Multimodal Task Specifications
- Baichuan 2: Open Large-scale Language Models
- CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning
- Evaluation and Analysis of Hallucination in Large Vision-Language Models
- Detecting and Preventing Hallucinations in Large Vision Language Models
- Image Synthesis under Limited Data: A Survey and Taxonomy
- BAGM: A Backdoor Attack for Manipulating Text-to-Image Generative Models
- RSGPT: A Remote Sensing Vision Language Model and Benchmark
- Med-Flamingo: a Multimodal Medical Few-shot Learner
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Lost in the Middle: How Language Models Use Long Contexts
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
- Grounded Text-to-Image Synthesis with Attention Refocusing
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
- NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation
- Evaluating Object Hallucination in Large Vision-Language Models
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models
- Qualitative Failures of Image Generation Models and Their Application in Detecting Deepfakes
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering
- Adding Conditional Control to Text-to-Image Diffusion Models
- Adding Conditional Control to Text-to-Image Diffusion Models
- Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Benchmarking Spatial Relationships in Text-to-Image Generation
- Scalable Diffusion Models with Transformers
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification
- InstructPix2Pix: Learning to Follow Image Editing Instructions
- Scaling Instruction-Finetuned Language Models
- Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training
- Flow Matching for Generative Modeling
- VIMA: General Robot Manipulation with Multimodal Prompts
- Brain Imaging Generation with Latent Diffusion Models
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object Detection
- Prompt-to-Prompt Image Editing with Cross Attention Control
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Flamingo: a Visual Language Model for Few-Shot Learning
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
- DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models
- Survey of Hallucination in Natural Language Generation
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Hallucination In Object Detection -- A Study In Visual Part Verification
- Learning Transferable Visual Models From Natural Language Supervision
- Zero-Shot Text-to-Image Generation
- SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
- Taming Transformers for High-Resolution Image Synthesis
- Center-based 3D Object Detection and Tracking
- Denoising Diffusion Probabilistic Models
- A Survey and Taxonomy of Adversarial Neural Networks for Text-to-Image Synthesis
- Multi-scale GANs for Memory-efficient Generation of High Resolution Medical Images
- BERTScore: Evaluating Text Generation with BERT
- Object Hallucination in Image Captioning
- Generative Image Inpainting with Contextual Attention
- Neural Discrete Representation Learning
- Deep reinforcement learning from human preferences
- An Analysis of Visual Question Answering Algorithms
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Microsoft COCO: Common Objects in Context
- Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
- LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
- Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
- T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
Cited by
Related