Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
2025/10/28 by Li, Yihan, Fu, Xiyuan, Verma, Ghanshyam +2 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.24476
Abstract
Hallucination remains one of the key obstacles to the reliable deployment of large language models (LLMs), particularly in real-world applications. Among various mitigation strategies, Retrieval-Augmented Generation (RAG) and reasoning enhancement have emerged as two of the most effective and widely adopted approaches, marking a shift from merely suppressing hallucinations to balancing creativity and reliability. However, their synergistic potential and underlying mechanisms for hallucination mitigation have not yet been systematically examined. This survey adopts an application-oriented perspective of capability enhancement to analyze how RAG, reasoning enhancement, and their integration in Agentic Systems mitigate hallucinations. We propose a taxonomy distinguishing knowledge-based and logic-based hallucinations, systematically examine how RAG and reasoning address each, and present a unified framework supported by real-world applications, evaluations, and benchmarks.
Citations
- A survey on retrieval-augmentation generation (RAG) models for healthcare applications
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit
- Agent-UniRAG: A Trainable Open-Source LLM Agent Framework for Unified Retrieval-Augmented Generation Systems
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- LLM4Ranking: An Easy-to-use Framework of Utilizing Large Language Models for Document Reranking
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
- XAttention: Block Sparse Attention with Antidiagonal Scoring
- Financial Analysis: Intelligent Financial Data Analysis System Based on LLM-RAG
- LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
- Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
- Medical Hallucinations in Foundation Models and Their Impact on Healthcare
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- KET-RAG: A Cost-Efficient Multi-Granular Indexing Framework for Graph-RAG
- The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
- s1: Simple test-time scaling
- CG-RAG: Research Question Answering by Citation Graph Retrieval-Augmented LLMs
- Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced Optimization
- WebWalker: Benchmarking LLMs in Web Traversal
- MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation
- SoK: Watermarking for AI-Generated Content
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- TART: An Open-Source Tool-Augmented Framework for Explainable Table-based Reasoning
- HALO: Hallucination Analysis and Learning Optimization to Empower LLMs with Retrieval-Augmented Context for Guided Clinical Decision Making
- LLMs Will Always Hallucinate, and We Need to Live With This
- LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain
- HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction
- Faculty Perspectives on the Potential of RAG in Computer Science Higher Education
- Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation
- AI models collapse when trained on recursively generated data
- GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
- Scaling Large Language Model-based Multi-Agent Collaboration
- RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation
- Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation
- Faithful Logical Reasoning via Symbolic Chain-of-Thought
- RaFe: Ranking Feedback Improves Query Rewriting for RAG
- A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- Hallucination of Multimodal Large Language Models: A Survey
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
- Is There No Such Thing as a Bad Question? H4R: HalluciBot For Ratiocination, Rewriting, Ranking, and Routing
- RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation
- LARA: Linguistic-Adaptive Retrieval-Augmentation for Multi-Turn Intent Classification
- RA-ISF: Learning to Answer and Understand from Retrieval Augmentation via Iterative Self-Feedback
- MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
- Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs
- SciAgent: Tool-augmented Language Models for Scientific Reasoning
- A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts
- A Survey on Large Language Model Hallucination via a Creativity Perspective
- Efficient Tool Use with Chain-of-Abstraction Reasoning
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
- Fine-grained Hallucination Detection and Editing for Language Models
- The Impact of Reasoning Step Length on Large Language Models
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
- Experiential Co-Learning of Software-Developing Agents
- Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models
- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
- Mistral 7B
- FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
- MKRAG: Medical Knowledge Retrieval Augmented Generation for Medical Question Answering
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- 🧜Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Code Llama: Open Foundation Models for Code
- An Empirical Study of the Non-Determinism of ChatGPT in Code Generation
- Reasoning in Large Language Models Through Symbolic Math Word Problems
- ChatDev: Communicative Agents for Software Development
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph
- Lost in the Middle: How Language Models Use Long Contexts
- BatGPT: A Bidirectional Autoregessive Talker from Generative Pre-trained Transformer
- WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences
- Deductive Verification of Chain-of-Thought Reasoning
- Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning
- Exposing Attention Glitches with Flip-Flop Language Modeling
- Query Rewriting for Retrieval-Augmented Large Language Models
- Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
- Structured Chain-of-Thought Prompting for Code Generation
- Structured Chain-of-Thought Prompting for Code Generation
- StarCoder: may the source be with you!
- Outline, Then Details: Syntactically Guided Coarse-To-Fine Code Generation
- The Internal State of an LLM Knows When It's Lying
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- REFINER: Reasoning Feedback on Intermediate Representations
- A Survey of Large Language Models
- Retrieving Multimodal Information for Augmented Generation: A Survey
- GPT-4 Technical Report
- A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
- LLaMA: Open and Efficient Foundation Language Models
- A Comprehensive Survey on Automatic Knowledge Graph Construction
- A Comprehensive Survey on Automatic Knowledge Graph Construction
- Toolformer: Language Models Can Teach Themselves to Use Tools
- In-Context Retrieval-Augmented Language Models
- Faithful Chain-of-Thought Reasoning
- Large Language Models Encode Clinical Knowledge
- Large language models encode clinical knowledge
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- PAL: Program-aided Language Models
- Contrastive Decoding: Open-ended Text Generation as Optimization
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
- Emergent Abilities of Large Language Models
- Factuality Enhanced Language Models for Open-Ended Text Generation
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Large Language Models are Zero-Shot Reasoners
- PaLM: Scaling Language Modeling with Pathways
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- Internet-augmented language models through few-shot prompting for open-domain question answering
- Survey of Hallucination in Natural Language Generation
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
- Fast Model Editing at Scale
- SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking
- Learning Passage Impacts for Inverted Indexes
- Editing Factual Knowledge in Language Models
- ProofWriter: Generating Implications, Proofs, and Abductive Statements\n over Natural Language
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- On Faithfulness and Factuality in Abstractive Summarization
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- Dense Passage Retrieval for Open-Domain Question Answering
- From Statistical Relational to Neuro-Symbolic Artificial Intelligence
- A Survey on Knowledge Graphs: Representation, Acquisition, and Applications
- GLTR: Statistical Detection and Visualization of Generated Text
- Attention Is All You Need
- Term-weighting approaches in automatic text retrieval
- TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related