Retrieval Augmentation Reduces Hallucination in Conversation
2021/04/15 by Kurt Shuster, Shuster, Kurt, Spencer Poff +7 · 149 citations
Computer Science · Psychology · #Artificial intelligence #Cognitive science #Communication #Computer science #Context (archaeology) #Conversation #Domain (mathematical analysis) #Encoder #Human–computer interaction #Natural Language Processing Techniques #Natural language processing #Psychology #Speech and dialogue systems #Task (project management) #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2104.07567
published in arXiv (Cornell University), 3784-3803 (Cornell University)
arxiv created 2021/04/15 · openalex publication_date 2021/04/15 · arxiv updated 2021/04/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Despite showing increasingly human-like conversational abilities, state-of-the-art dialogue models often suffer from factual incorrectness and hallucination of knowledge (Roller et al., 2020). In this work we explore the use of neural-retrieval-in-the-loop architectures - recently shown to be effective in open-domain QA (Lewis et al., 2020b; Izacard and Grave, 2020) - for knowledge-grounded dialogue, a task that is arguably more challenging as it requires querying based on complex multi-turn dialogue context and generating conversationally coherent responses. We study various types of architectures with multiple components - retrievers, rankers, and encoder-decoders - with the goal of maximizing knowledgeability while retaining conversational ability. We demonstrate that our best models obtain state-of-the-art performance on two knowledge-grounded conversational tasks. The models exhibit open-domain conversational capabilities, generalize effectively to scenarios not within the training data, and, as verified by human evaluations, substantially reduce the well-known problem of knowledge hallucination in state-of-the-art chatbots.
Cited by
- Beyond the Prompt: An Empirical Study of Cursor Rules
- DACE For Railway Acronym Disambiguation
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- FloodSQL-Bench: A Retrieval-Augmented Benchmark for Geospatially-Grounded Text-to-SQL
- Adjudicator: Correcting Noisy Labels with a KG-Informed Council of LLM Agents
- Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering
- A Systematic Framework for Enterprise Knowledge Retrieval: Leveraging LLM-Generated Metadata to Enhance RAG Systems
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment
- A Concise Review of Hallucinations in LLMs and their Mitigation
- Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%
- Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
- Measuring the Impact of Lexical Training Data Coverage on Hallucination Detection in Large Language Models
- Beyond Component Strength: Synergistic Integration and Adaptive Calibration in Multi-Agent RAG Systems
- RAGSmith: A Framework for Finding the Optimal Composition of Retrieval-Augmented Generation Methods Across Datasets
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
- CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
- Catching Contamination Before Generation: Spectral Kill Switches for Agents
- COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
- ContextPilot: Fast Long-Context Inference via Context Reuse
- Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- EncouRAGe: Evaluating RAG Local, Fast, and Reliable
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- MedFusionT5: Cross-Modal Attention Boosts Semantic Quality and Reduces Hallucinations in Dental AI
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- PGDA-KGQA: A Prompt-Guided Generative Framework with Multiple Data Augmentation Strategies for Knowledge Graph Question Answering
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- Learning Efficient and Generalizable Graph Retriever for Knowledge-Graph Question Answering
- Interpretability Framework for LLMs in Undergraduate Calculus
- Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
- RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge
- Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- VeriCite: Towards Reliable Citations in Retrieval-Augmented Generation via Rigorous Verification
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
- FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
- LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- MimiTalk: Revolutionizing Qualitative Research with Dual-Agent AI
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
- Enhancing LLM-based Fault Localization with a Functionality-Aware Retrieval-Augmented Generation Framework
- CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
- ReGeS: Reciprocal Retrieval-Generation Synergy for Conversational Recommender Systems
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- Learning the natural history of human disease with generative transformers
- HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
- ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
- Large Language Models Meet Legal Artificial Intelligence: A Survey
- DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling
- Fact or Facsimile? Evaluating the Factual Robustness of Modern Retrievers
- From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
- AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark
- Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
- Hallucinations in medical devices
- Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
- Classification is a RAG problem: A case study on hate speech detection
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- A Pragmatist Robot: Learning to Plan Tasks by Experiencing the Real World
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- ASINT: Learning AS-to-Organization Mapping from Internet Metadata
- Simple Methods Defend RAG Systems Well Against Real-World Attacks
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- MUST-RAG: MUSical Text Question Answering with Retrieval Augmented Generation
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation
- A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
- Hallucination Detection and Mitigation with Diffusion in Multi-Variate Time-Series Foundation Models
- DeepResearchEco: A Recursive Agentic Workflow for Complex Scientific Question Answering in Ecology
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- Grahak-Nyay: Consumer Grievance Redressal through Large Language Models
- Dynamic Injection of Entity Knowledge into Dense Retrievers
- KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
- RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
- Machine Assistant with Reliable Knowledge: Enhancing Student Learning via RAG-based Retrieval
- What Characteristics Make ChatGPT Effective for Software Issue Resolution? An Empirical Study of Task, Project, and Conversational Signals in GitHub Issues
- Retrieval-Confused Generation is a Good Defender for Privacy Violation Attack of Large Language Models
- Controlling Context: Generative AI at Work in Integrated Circuit Design and Other High-Precision Domains
- Re-Initialization Token Learning for Tool-Augmented Large Language Models
- SlimRAG: Retrieval without Graphs via Entity-Aware Context Selection
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
- Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases
- Large Language Model-Powered Conversational Agent Delivering Problem-Solving Therapy (PST) for Family Caregivers: Enhancing Empathy and Therapeutic Alliance Using In-Context Learning
- Respecting Temporal-Causal Consistency: Entity-Event Knowledge Graphs for Retrieval-Augmented Generation
- Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning
- CLATTER: Comprehensive Entailment Reasoning for Hallucination Detection
- Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
- Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- Retrieval Augmented Time Series Forecasting
- Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency
- SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL
- An Integrated Platform for LEED Certification Automation Using Computer Vision and LLM-RAG
- MEGA-GPT: Artificial Intelligence Guidance and Building Analytical Protocols Using MEGA Software
- MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
- Safeguarding Privacy of Retrieval Data against Membership Inference Attacks: Is This Query Too Close to Home?
- Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
- Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
- TeroSeek: An AI-Powered Knowledge Base and Retrieval Generation Platform for Terpenoid Research
- Prompting is not Enough: Exploring Knowledge Integration and Controllable Generation
- syftr: Pareto-Optimal Generative AI
- Knoll: Creating a Knowledge Ecosystem for Large Language Models
- LLMs for Supply Chain Management
- Evidence-Grounded Multimodal Misinformation Detection with Attention-Based GNNs
- LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
- CUB: Benchmarking Context Utilisation Techniques for Language Models
- The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
- Safety Degradation in AI Agents
- Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
- Transparent and Robust RAG: Adaptive-Reward Reinforcement Learning for Decision Traceability
- Reasoning BO: Enhancing Bayesian Optimization with Long-Context Reasoning Power of LLMs
- GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs
- RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models
- IterKey: Iterative Keyword Generation with LLMs for Enhanced Retrieval Augmented Generation
- AI-Based Speaking Assistant: Supporting Non-Native Speakers' Speaking in Real-Time Multilingual Communication
- RASPRef: Retrieval-Augmented Self-Supervised Prompt Refinement for Large Reasoning Models
- EnronQA: Towards Personalized RAG over Private Documents
- IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
- Detecting Manipulated Contents Using Knowledge-Grounded Inference
- ImproBR: Bug Report Improver Using LLMs
- Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
- Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
- Conflicts in Texts: Data, Implications and Challenges
- Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
- RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
- RPO-RAG: Aligning Small LLMs with Relation-aware Preference Optimization for Knowledge Graph Question Answering
- Bridging Instead of Replacing Online Coding Communities with AI through Community-Enriched Chatbot Designs
- RAG-Stack: Co-Optimizing RAG Serving Performance and Quality
- Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
- CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction
- WindVE: Collaborative CPU-NPU Vector Embedding
- LegalRAG: A Hybrid RAG System for Multilingual Legal Information Retrieval
- Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
- Out of Style: RAG's Fragility to Linguistic Variation
- Plan-and-Refine: Diverse and Comprehensive Retrieval-Augmented Generation
- HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
- Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey
- Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
Related