A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
2023/11/09 by Lei Huang, Weijiang Yu, Weitao Ma +8 · 4 voices · 421 citations
Computer Science · Engineering · #Ferroelectric and Negative Capacitance Devices #Text Readability and Simplification #Topic Modeling #cs.CL
paper · pdf · doi:10.1145/3703155
openalex publication_date 2024/11/20 · openalex created_date 2024/11/21 · openalex updated_date 2026/07/31
Abstract
The emergence of large language models (LLMs) has marked a significant breakthrough in natural language processing (NLP), fueling a paradigm shift in information acquisition. Nevertheless, LLMs are prone to hallucination, generating plausible yet nonfactual content. This phenomenon raises significant concerns over the reliability of LLMs in real-world information retrieval (IR) systems and has attracted intensive research to detect and mitigate such hallucinations. Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.
Citations
Cited by
- Investigating Users' Search Behavior and Outcome with ChatGPT in Learning-oriented Search Tasks
- What large language models know and what people think they know
- From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
- A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs
- Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
- Epistemic Familiarity is Associated With Belief Stability in Large Language Models
- Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
- DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
- It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation
- The unintended consequences of large language models as a labor-augmenting technology in science
- AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization
- DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening
- OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
- Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
- Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
- Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
- The CRAFT principles for the responsible use of large language models in policymaking
- Is External Database Protection Static in Retrieval-Augmented Generation? Rethinking Privacy Preservation under Dynamic Queries
- MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA
- Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge Graphs
- Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Accurate and Efficient Long-Term Memory for LLM Agents
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- MIRAGE: The Illusion of Visual Understanding
- Training LLMs for Honesty via Confessions
- Critical Confabulation: Can LLMs Hallucinate for Social Good?
- Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0
- On Non-interactive Evaluation of Animal Communication Translators
- Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- Whither symbols in the era of advanced neural networks?
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
- Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Connectomics Informed by Large Language Models
- On Conditional Stochastic Interpolation for Generative Nonlinear Sufficient Dimension Reduction
- From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change
- Boosting metacognition in entangled human-AI interaction to navigate cognitive-behavioral drift
- Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
- Towards end-to-end automation of AI research
- Automated extraction of toxicological mechanism of action from PubMed literature using Large Language Models
- With Visual Integrity and Care: A Framework for Mixed Methods Research on Visual Social Data
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
- Expert-Grounded Automatic Prompt Engineering for Extracting Lattice Constants of High-Entropy Alloys from Scientific Publications using Large Language Models
- TARGET: Automated Scenario Generation from Traffic Rules for Testing Autonomous Vehicles via Validated LLM-Guided Knowledge Extraction
- Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation
- Hallucination Rates in Language Generation
- Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
- TRE: Training-Free Hallucination Detection for Diffusion Language Models
- SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
- A Unified Definition of Hallucination, Or: It's the World Model, Stupid
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Event Extraction in Large Language Model
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- Eidoku: A Neuro-Symbolic Verification Gate for LLM Reasoning via Structural Constraint Satisfaction
- CheXPO-v2: Preference Optimization for Chest X-ray VLMs with Knowledge Graph Consistency
- Algorithmic UDAP
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- Collaborative Edge-to-Server Inference for Vision-Language Models
- Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting
- Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering
- Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
- AIAuditTrack: A Framework for AI Security system
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
- Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC
- FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
- dtreg: Describing Data Analysis in Machine-Readable Format in Python and R
- Phythesis: Physics-Guided Evolutionary Scene Synthesis for Energy-Efficient Data Center Design via LLMs
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- LLMs in Interpreting Legal Documents
- Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines
- Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
- CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models
- Calibrated Trust in Dealing with LLM Hallucinations: A Qualitative Study
- Multi-Agent Intelligence for Multidisciplinary Decision-Making in Gastrointestinal Oncology
- Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- Language-driven Fine-grained Retrieval
- Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platforms
- Trusted AI Agents in the Cloud
- LLM Harms: A Taxonomy and Discussion
- The Road of Adaptive AI for Precision in Cybersecurity
- Latent Debate: A Surrogate Framework for Interpreting LLM Thinking
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- The Personalization Paradox: Semantic Loss vs. Reasoning Gains in Agentic AI Q&A
- A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
- Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study
- Table as a Modality for Large Language Models
- Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
- Toward a Safe Internet of Agents
- MATCH: Engineering Transparent and Controllable Conversational XAI Systems through Composable Building Blocks
- Factors That Support Grounded Responses in LLM Conversations: A Rapid Review
- REFLEX: Self-Refining Explainable Fact-Checking via Disentangling Truth into Style and Substance
- Enhancing Sequential Recommendation with World Knowledge from Large Language Models
- HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization
- HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
- Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
- What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
- Towards Automating Data Access Permissions in AI Agents
- Synthesizing Precise Protocol Specs from Natural Language for Effective Test Generation
- An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
- Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning
- Build AI Assistants using Large Language Models and Agents to Enhance the Engineering Education of Biomechanics
- Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- Can large language models be a cardinality estimator? An empirical study
- BeautyGuard: Designing a Multi-Agent Roundtable System for Proactive Beauty Tech Compliance through Stakeholder Collaboration
- Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
- Are LLMs The Way Forward? A Case Study on LLM-Guided Reinforcement Learning for Decentralized Autonomous Driving
- On the Influence of Artificial Intelligence on Human Problem-Solving: Empirical Insights for the Third Wave in a Multinational Longitudinal Pilot Study
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- A multimodal AI agent for clinical decision support in ophthalmology
- Unreliable minds, unreliable machines: dyslexic memory, ChatGPT, and the epistemic disobedience of generative AI
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- Trading Vector Data in Vector Databases
- Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
- A Low-Rank Method for Vision Language Model Hallucination Mitigation in Autonomous Driving
- AIA Forecaster: Technical Report
- Evaluation of retrieval-based QA on QUEST-LOFT
- SymLight: Exploring Interpretable and Deployable Symbolic Policies for Traffic Signal Control
- AI-assisted workflow enables rapid, high-fidelity breast cancer clinical trial eligibility prescreening
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- ORCHID: Orchestrated Retrieval-Augmented Classification with Human-in-the-Loop Intelligent Decision-Making for High-Risk Property
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
- KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering
- LLM-enhanced Air Quality Monitoring Interface via Model Context Protocol
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- Using language models to label clusters of scientific documents
- TreeQA: Enhanced LLM-RAG with logic tree reasoning for reliable and interpretable multi-hop question answering
- Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
- A Dual Large Language Models Architecture with Herald Guided Prompts for Parallel Fine Grained Traffic Signal Control
- Retrieval Augmented Generation-Enhanced Distributed LLM Agents for Generalizable Traffic Signal Control with Emergency Vehicles
- Structurally Valid Log Generation using FSM-GFlowNets
- E-Scores for (In)Correctness Assessment of Generative Model Outputs
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- Position: Biology is the Challenge Physics-Informed ML Needs to Evolve
- AnnoBench: A Benchmark for Visualization Annotation Generation
- Machine learning in biological research: key algorithms, applications, and future directions
- From limitation to possibility: affordances and tactical evolution in user–GenAI interaction
- How to Use Generative AI in Educational Research
- Learning from Wildfire Decision Support: large language model analysis of barriers to fire spread in a census of large wildfires in the United States (2011–2023)
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
- BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Hallucination Benchmark for Speech Foundation Models
- Bowling with ChatGPT: On the Evolving User Interactions with Conversational AI Systems
- How well LLM-based test generation techniques perform with newer LLM versions?
- "I Use ChatGPT to Humanize My Words": Affordances and Risks of ChatGPT to Autistic Users
- Mapping and Comparing Climate Equity Policy Practices Using RAG LLM-Based Semantic Analysis and Recommendation Systems
- Ontology-driven generation of parameters for health technology assessment models: a prompt engineering study
- Large-scale manual curation and harmonization of metadata from metagenomic and cancer genomic repositories: challenges and solutions
- What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
- Hallucination Localization in Video Captioning
- Model-Document Protocol for AI Search
- PICOs-RAG: PICO-supported Query Rewriting for Retrieval-Augmented Generation in Evidence-Based Medicine
- M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
- Estimating the Error of Large Language Models at Pairwise Text Comparison
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- InterpDetect: Interpretable Signals for Detecting Hallucinations in Retrieval-Augmented Generation
- DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical services
- MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs Fuzzing
- SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation
- Neural Diversity Regularizes Hallucinations in Small Models
- HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
- FreeChunker: A Cross-Granularity Chunking Framework
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- Enhancing visual-LLM for construction site safety compliance via prompt engineering and Bi-stage retrieval-augmented generation
- LLM-Augmented Symbolic NLU System for More Reliable Continuous Causal Statement Interpretation
- A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
- ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
- The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
- Beyond "Hallucinations": A Framework for Stable Human-AI Reasoning
- Automated Extraction of Protocol State Machines from 3GPP Specifications with Domain-Informed Prompts and LLM Ensembles
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Element2Vec: Build Chemical Element Representation from Text for Property Prediction
- Counting Hallucinations in Diffusion Models
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG
- ESI: Epistemic Uncertainty Quantification via Semantic-preserving Intervention for Large Language Models
- Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for Recommendation
- Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
- iCodeReviewer: Improving Secure Code Review with Mixture of Prompts
- Credal Transformer: A Principled Approach for Quantifying and Mitigating Hallucinations in Large Language Models
- Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions
- AgentCaster: Reasoning-Guided Tornado Forecasting
- Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
- Learning to Reason for Hallucination Span Detection
- From Craft to Constitution: A Governance-First Paradigm for Principled Agent Engineering
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- Proof Strategy Extraction from LLMs for Enhancing Symbolic Provers
- CacheClip: Accelerating RAG with Effective KV Cache Reuse
- Failure-Driven Workflow Refinement
- The trust crisis in artificial intelligence: AI hallucinations and human-AI collaboration
- Artificial intelligence in web accessibility: A systematic mapping study
- Speed Kills: Exploring Confused Deputy Attacks Through Edge AI Accelerators
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Fundamentals of Building Autonomous LLM Agents
- Man-Made Heuristics Are Dead. Long Live Code Generators!
- Revisiting Hallucination Detection with Effective Rank-based Uncertainty
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- Stop DDoS Attacking the Research Community with AI-Generated Survey Papers
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Towards Reliable Retrieval in RAG Systems for Large Legal Datasets
- Artificial Intelligence in Detecting Statistical Errors: Implications for Authors, Reviewers, and Editors
- FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
- AgentDR Dynamic Recommendation with Implicit Item-Item Relations via LLM-based Agents
- Domain-Shift-Aware Conformal Prediction for Large Language Models
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- A novel hallucination classification framework
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
- Large Language Models Hallucination: A Comprehensive Survey
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Evaluating Large Language Models for IUCN Red List Species Information
- Geolog-IA: Conversational System for Academic Theses
- Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
- Energy-Regularized Sequential Model Editing on Hyperspheres
- Data Quality Challenges in Retrieval-Augmented Generation
- Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
- Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- Agentic Exploration of Physics Models
- Hallucination is Inevitable for LLMs with the Open World Assumption
- Neural Message-Passing on Attention Graphs for Hallucination Detection
- Adoption paradox of artificial intelligence in computational pathology: a three-stage maturity model from algorithms to clinical integration
- Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
- CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
- Small Language Models for Curriculum-based Guidance
- What If Moderation Didn't Mean Suppression? A Case for Personalized Content Transformation
- Smoothing-Based Conformal Prediction for Balancing Efficiency and Interpretability
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- A model of errors in transformers
- LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
- Are Hallucinations Bad Estimations?
- Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- Generative AI for FFRDCs
- OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
- Augmented and Programmatically Optimized LLM Prompts Reduce Chemical Hallucinations
- Retrieval-Augmented Generation – Ein Erfahrungsbericht und Leitfaden zum Einsatz von wissensbasierten Chatbots an der HU Berlin (Teil 1)
- Evaluating large language models for accuracy incentivizes hallucinations
- How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
- Reflection and Experimental Rigor Are Our AiMS: A New Metacognitive Framework for Experimental Design
- Identifying and Addressing User-level Security Concerns in Smart Homes Using "Smaller" LLMs
- A Knowledge Graph and a Tripartite Evaluation Framework Make Retrieval-Augmented Generation Scalable and Transparent
- Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
- CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
- EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol
- Campus AI vs. Commercial AI: How Customizations Shape Trust and Usage of LLM as-a-Service Chatbots
- Toward Efficient Influence Function: Dropout as a Compression Tool
- Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- SynBench: A Benchmark for Differentially Private Text Generation
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- SCORE: A Semantic Evaluation Framework for Generative Document Parsing
- Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs
- SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
- Root Cause Analysis of Radiation Oncology Incidents Using Large Language Models
- KoSEL: Knowledge subgraph enhanced large language model for medical question answering
- MMORE: Massive Multimodal Open RAG & Extraction
- HARP: Hallucination Detection via Reasoning Subspace Projection
- Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking
- Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability
- Large Language Models for Security Operations Centers: A Comprehensive Survey
- Virtual Agent Economies
- Robo-Advisors Beyond Automation: Principles and Roadmap for AI-Driven Financial Planning
- Robot guide with multi-agent control and automatic scenario generation with LLM
- Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- Changing the Paradigm from Dynamic Queries to LLM-generated SQL Queries with Human Intervention
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems
- LightAgent: Production-level Open-source Agentic AI Framework
- Deploying AI for Signal Processing education: Selected challenges and intriguing opportunities
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Trees
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
- Temporal Counterfactual Explanations of Behaviour Tree Decisions
- AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
- Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
- Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
- KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino
- Stack Overflow Is Not Dead Yet: Crowd Answers Still Matter
- Prior Distribution and Model Confidence
- Conversational AI increases political knowledge as effectively as self-directed internet search
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Can LLMs Lie? Investigation beyond Hallucination
- DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
- An LLM-enabled semantic-centric framework to consume privacy policies
- Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- Exploring and Mitigating Fawning Hallucinations in Large Language Models
- TECP: Token-Entropy Conformal Prediction for LLMs
- BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting
- ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Data Auctions for Retrieval Augmented Generation
- Data-driven Discovery of Digital Twins in Biomedical Research
- Multi-Ontology Integration with Dual-Axis Propagation for Medical Concept Representation
- Addressing accuracy and hallucination of LLMs in Alzheimer's disease research through knowledge graphs
- From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
- Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems
- Real-Time Detection of Hallucinated Entities in Long-Form Generation
- A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants
- ConfTuner: Training Large Language Models to Express Their Confidence Verbally
- Skeptik: A Hybrid Framework for Combating Potential Misinformation in Journalism
- Principled Detection of Hallucinations in Large Language Models via Multiple Testing
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Preguss: It Analyzes, It Specifies, It Verifies
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Driving Style Recognition Like an Expert Using Semantic Privileged Information from Large Language Models
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- Applied causality to infer protein dynamics and kinetics
- Can we Evaluate RAGs with Synthetic Data?
- Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- Never compromise with vulnerabilities: a comprehensive survey on AI governance
- MuaLLM: A Multimodal Large Language Model Agent for Circuit Design Assistance with Hybrid Contextual Retrieval-Augmented Generation
- Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
- Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
- Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials
- LinguaFluid: Language Guided Fluid Control via Semantic Rewards in Reinforcement Learning
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction Modelling
- StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback
- Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
- LLM-Prior: A Framework for Knowledge-Driven Prior Elicitation and Aggregation
- ASINT: Learning AS-to-Organization Mapping from Internet Metadata
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
- MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
- L3M+P: Lifelong Planning with Large Language Models
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
- HySemRAG: A Hybrid Semantic Retrieval-Augmented Generation Framework for Automated Literature Synthesis and Methodological Gap Analysis
- DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- Latent Knowledge Scalpel: Precise and Massive Knowledge Editing for Large Language Models
- Audio Prototypical Network For Controllable Music Recommendation
- Accessibility Scout: Personalized Accessibility Scans of Built Environments
- How Far Are AI Scientists from Changing the World?
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
Discussions
- Quellen: dl.acm.org/doi/full/10...., arxiv.org/html/2509.22..., help.openai.com/en/articles/... [bsky, 17 points, 3 comments]
- Spannendes Uni-Projekt. KI-Hallizunationen: Ursachen, Wirkung, Lösungen. Ich bin Team "Ursachen". Das Paper schlägt eine strukturierte Taxonomie vor: Es gliedert Halluzinationen in Kategorien, um Klar [bsky, 9 points, 1 comments]
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions https://arxiv.org/abs/2311.05232v1 #AI #hallucination [bsky, 1 points, 0 comments]
- arxiv.org/abs/2311.05232 [bsky, 0 points, 1 comments]
Related