On Faithfulness and Factuality in Abstractive Summarization
2020/05/02 by Joshua Maynez, Maynez, Joshua, Shashi Narayan +5 · 142 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
paper · pdf · doi:10.48550/arxiv.2005.00661
ACL 2020, 14 pages
arxiv created 2020/05/02 · arxiv updated 2020/05/05
Abstract
It is well known that the standard likelihood training and approximate decoding objectives in neural text generation models lead to less human-like responses for open-ended tasks such as language modeling and story generation. In this paper we have analyzed limitations of these models for abstractive document summarization and found that these models are highly prone to hallucinate content that is unfaithful to the input document. We conducted a large scale human evaluation of several neural abstractive summarization systems to better understand the types of hallucinations they produce. Our human annotators found substantial amounts of hallucinated content in all model generated summaries. However, our analysis does show that pretrained models are better summarizers not only in terms of raw metrics, i.e., ROUGE, but also in generating faithful and factual summaries as evaluated by humans. Furthermore, we show that textual entailment measures better correlate with faithfulness than standard metrics, potentially leading the way to automatic evaluation metrics as well as training and decoding criteria.
Cited by
- ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- TRE: Training-Free Hallucination Detection for Diffusion Language Models
- Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL
- A Unified Definition of Hallucination: It's The World Model, Stupid!
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation
- Ev-Trust: An Evolutionarily Stable Trust Mechanism for Decentralized LLM-Based Multi-Agent Service Economies
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Q-BAR: Blogger Anomaly Recognition via Quantum-enhanced Manifold Learning
- Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- Learning from Self Critique and Refinement for Faithful LLM Summarization
- Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%
- Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
- "AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
- Can large language models be a cardinality estimator? An empirical study
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- Trustworthy LLM-Mediated Communication: Evaluating Information Fidelity in LLM as a Communicator (LAAC) Framework in Multiple Application Domains
- Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
- Integrating Ontologies with Large Language Models for Enhanced Control Systems in Chemical Engineering
- The Geometry of Dialogue: Graphing Language Models to Reveal Synergistic Teams for Multi-Agent Collaboration
- TextualVerifier: Verify TextGrad Step-by-Step
- Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
- OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
- Reward Models are Metrics in a Trench Coat
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- MedFusionT5: Cross-Modal Attention Boosts Semantic Quality and Reduces Hallucinations in Dental AI
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- Harnessing the Power of Large Language Models for Software Testing Education: A Focus on ISTQB Syllabus
- Neural Diversity Regularizes Hallucinations in Language Models
- Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- Gobernanza y trazabilidad "a prueba de AI Act" para casos de uso legales: un marco técnico-jurídico, métricas forenses y evidencias auditables
- Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
- Enhancing Faithfulness in Abstractive Summarization via Span-Level Fine-Tuning
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Text2Stories: Evaluating the Alignment Between Stakeholder Interviews and Generated User Stories
- Domain-Shift-Aware Conformal Prediction for Large Language Models
- InforME: Improving Informativeness of Abstractive Text Summarization With Informative Attention Guided by Named Entity Salience
- Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
- Large Language Models Hallucination: A Comprehensive Survey
- Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- Public Opinion on the Politics of AI Alignment: Cross-National Evidence on Expectations for AI Moderation From Germany and the United States
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
- SUIT: Knowledge Editing with Subspace-Aware Key-Value Mappings
- Machines in the Margins: A Systematic Review of Automated Content Generation for Wikipedia
- PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
- InsightGUIDE: An Opinionated AI Assistant for Guided Critical Reading of Scientific Literature
- Retrieval Feedback Memory Enhancement Large Model Retrieval Generation Method
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- A multi-query, multimodal, receiver-augmented solution to extract contemporary cardiology guideline information using large language models
- How Large Language Models are Designed to Hallucinate
- A funny companion: Distinct neural responses to AI- versus human-attributed humor
- HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
- MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems
- Deploying AI for Signal Processing education: Selected challenges and intriguing opportunities
- Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection
- From Confidence to Collapse in LLM Factual Robustness
- Grounding the Ungrounded: A Spectral-Graph Framework for Quantifying Hallucinations in Multimodal LLMs
- LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
- CUTE-MRI: Conformalized Uncertainty-based framework for Time-adaptivE MRI
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
- Hallucination Detection and Mitigation in Scientific Text Simplification using Ensemble Approaches: DS@GT at CLEF 2025 SimpleText
- Can we Evaluate RAGs with Synthetic Data?
- Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
- Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
- Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- ChartCap: Mitigating Hallucination of Dense Chart Captioning
- Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- Investigating Hallucination in Conversations for Low Resource Languages
- Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
- Theoretical Foundations and Mitigation of Hallucination in Large Language Models
- Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models
- A Mathematical Theory of Discursive Networks
- Spatial Visual Analytics for Multi-Document Summary Verification
- REAL: Benchmarking Abilities of Large Language Models for Housing Transactions and Services
- Who's Sorry Now: User Preferences Among Rote, Empathic, and Explanatory Apologies from LLM Chatbots
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Response Quality Assessment for Retrieval-Augmented Generation via Conditional Conformal Factuality
- Cosmos: Compressed and Smooth Latent Space for Text Diffusion Modeling
- Generative AI without guardrails can harm learning: Evidence from high school mathematics
- HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations
- TermSight: Making Service Contracts Approachable
- ConfRAG: Confidence-Guided Retrieval-Augmenting Generation
- Textual Bayes: Quantifying Uncertainty in LLM-Based Systems
- Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know'
- Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
- Reviewriter: AI-Generated Instructions For Peer Review Writing
- World Modelling Improves Language Model Agents
- MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
- Fact-Controlled Diagnosis of Hallucinations in Medical Text Summarization
- Interpretable Zero-shot Learning with Infinite Class Concepts
- Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
- TRAPDOC: Deceiving LLM Users by Injecting Imperceptible Phantom Tokens into Documents
- Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model
- Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis
- AITEE -- Agentic Tutor for Electrical Engineering
- Pretrained LLMs Learn Multiple Types of Uncertainty
- Between fact and fairy: tracing the hallucination metaphor in AI discourse
- Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
- The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
- Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs
- LLM-Powered AI Agent Systems and Their Applications in Industry
- Aug2Search: Enhancing Facebook Marketplace Search with LLM-Generated Synthetic Data Augmentation
- Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
- When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification
- Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
- A Survey on Foundation Models for Personalized Federated Intelligence
- Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
- ”My AI is Lying to Me”: User-reported LLM hallucinations in AI mobile apps reviews
- Artificial Intelligence in Government: Why People Feel They Lose Control
- Multi-agents based User Values Mining for Recommendation
- Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines
- Consistency in Language Models: Current Landscape, Challenges, and Future Directions
- Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity
- Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
- Information Retrieval in the Age of Generative AI: The RGB Model
- PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
- Towards Long Context Hallucination Detection
- Random-Set Large Language Models
- Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding
- ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
- Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression
- Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
- ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
- Explainable AI in Usable Privacy and Security: Challenges and Opportunities
- Purposefully Induced Psychosis (PIP): Embracing Hallucination as Imagination in Large Language Models
- Cross-platform epistemic verification for improving factual reliability in AI-generated news summarization
- HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
- Large language models as uncertainty-calibrated optimizers for experimental discovery
Related