Retrieval Augmentation Reduces Hallucination in Conversation
2021/04/15 by Kurt Shuster, Shuster, Kurt, Spencer Poff +7 · 73 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2104.07567
Abstract
Despite showing increasingly human-like conversational abilities, state-of-the-art dialogue models often suffer from factual incorrectness and hallucination of knowledge (Roller et al., 2020). In this work we explore the use of neural-retrieval-in-the-loop architectures - recently shown to be effective in open-domain QA (Lewis et al., 2020b; Izacard and Grave, 2020) - for knowledge-grounded dialogue, a task that is arguably more challenging as it requires querying based on complex multi-turn dialogue context and generating conversationally coherent responses. We study various types of architectures with multiple components - retrievers, rankers, and encoder-decoders - with the goal of maximizing knowledgeability while retaining conversational ability. We demonstrate that our best models obtain state-of-the-art performance on two knowledge-grounded conversational tasks. The models exhibit open-domain conversational capabilities, generalize effectively to scenarios not within the training data, and, as verified by human evaluations, substantially reduce the well-known problem of knowledge hallucination in state-of-the-art chatbots.
Cited by
- Beyond the Prompt: An Empirical Study of Cursor Rules
- DACE For Railway Acronym Disambiguation
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- FloodSQL-Bench: A Retrieval-Augmented Benchmark for Geospatially-Grounded Text-to-SQL
- Adjudicator: Correcting Noisy Labels with a KG-Informed Council of LLM Agents
- Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering
- A Systematic Framework for Enterprise Knowledge Retrieval: Leveraging LLM-Generated Metadata to Enhance RAG Systems
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment
- A Concise Review of Hallucinations in LLMs and their Mitigation
- Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%
- Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
- Measuring the Impact of Lexical Training Data Coverage on Hallucination Detection in Large Language Models
- Beyond Component Strength: Synergistic Integration and Adaptive Calibration in Multi-Agent RAG Systems
- RAGSmith: A Framework for Finding the Optimal Composition of Retrieval-Augmented Generation Methods Across Datasets
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
- CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
- Catching Contamination Before Generation: Spectral Kill Switches for Agents
- COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
- ContextPilot: Fast Long-Context Inference via Context Reuse
- Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- EncouRAGe: Evaluating RAG Local, Fast, and Reliable
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- MedFusionT5: Cross-Modal Attention Boosts Semantic Quality and Reduces Hallucinations in Dental AI
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- Interpretability Framework for LLMs in Undergraduate Calculus
- Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
- RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge
- Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- VeriCite: Towards Reliable Citations in Retrieval-Augmented Generation via Rigorous Verification
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
- FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
- LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- MimiTalk: Revolutionizing Qualitative Research with Dual-Agent AI
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
- Enhancing LLM-based Fault Localization with a Functionality-Aware Retrieval-Augmented Generation Framework
- CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
- ReGeS: Reciprocal Retrieval-Generation Synergy for Conversational Recommender Systems
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- Learning the natural history of human disease with generative transformers
- HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
- ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
- Large Language Models Meet Legal Artificial Intelligence: A Survey
- DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling
- Fact or Facsimile? Evaluating the Factual Robustness of Modern Retrievers
- From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
- AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark
- Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
- Hallucinations in medical devices
- Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
- Classification is a RAG problem: A case study on hate speech detection
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- ASINT: Learning AS-to-Organization Mapping from Internet Metadata
- Simple Methods Defend RAG Systems Well Against Real-World Attacks
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- MUST-RAG: MUSical Text Question Answering with Retrieval Augmented Generation
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation
Related