A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
2023/11/09 by Lei Huang, Weijiang Yu, Weitao Ma +8 · 4 voices · 1,757 citations
Computer Science · Engineering · Psychology · #Cognitive psychology #Ferroelectric and Negative Capacitance Devices #Psychiatry #Psychology #Text Readability and Simplification #Topic Modeling #Visual Hallucination #cs.CL
paper · pdf · doi:10.1145/3703155
published in ACM Transactions on Information Systems 43(2), 1-55
openalex publication_date 2024/11/20 · openalex created_date 2024/11/21 · openalex updated_date 2026/08/05
Abstract
The emergence of large language models (LLMs) has marked a significant breakthrough in natural language processing (NLP), fueling a paradigm shift in information acquisition. Nevertheless, LLMs are prone to hallucination, generating plausible yet nonfactual content. This phenomenon raises significant concerns over the reliability of LLMs in real-world information retrieval (IR) systems and has attracted intensive research to detect and mitigate such hallucinations. Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.
Citations
Cited by
- Investigating Users' Search Behavior and Outcome with ChatGPT in Learning-oriented Search Tasks
- What large language models know and what people think they know
- From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
- A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs
- Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
- Epistemic Familiarity is Associated With Belief Stability in Large Language Models
- Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
- DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
- It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation
- The unintended consequences of large language models as a labor-augmenting technology in science
- AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization
- DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening
- OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
- Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
- Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
- Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
- The CRAFT principles for the responsible use of large language models in policymaking
- Is External Database Protection Static in Retrieval-Augmented Generation? Rethinking Privacy Preservation under Dynamic Queries
- MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA
- Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge Graphs
- Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Accurate and Efficient Long-Term Memory for LLM Agents
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- MIRAGE: The Illusion of Visual Understanding
- Training LLMs for Honesty via Confessions
- Critical Confabulation: Can LLMs Hallucinate for Social Good?
- Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0
- On Non-interactive Evaluation of Animal Communication Translators
- Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- Whither symbols in the era of advanced neural networks?
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
- Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Connectomics Informed by Large Language Models
- On Conditional Stochastic Interpolation for Generative Nonlinear Sufficient Dimension Reduction
- From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change
- Boosting metacognition in entangled human-AI interaction to navigate cognitive-behavioral drift
- Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
- Towards end-to-end automation of AI research
- Automated extraction of toxicological mechanism of action from PubMed literature using Large Language Models
- With Visual Integrity and Care: A Framework for Mixed Methods Research on Visual Social Data
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
- Expert-Grounded Automatic Prompt Engineering for Extracting Lattice Constants of High-Entropy Alloys from Scientific Publications using Large Language Models
- TARGET: Automated Scenario Generation from Traffic Rules for Testing Autonomous Vehicles via Validated LLM-Guided Knowledge Extraction
- Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation
- Hallucination Rates in Language Generation
- Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
- TRE: Training-Free Hallucination Detection for Diffusion Language Models
- SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
- A Unified Definition of Hallucination: It's The World Model, Stupid!
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Event Extraction in Large Language Model
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- Eidoku: A Neuro-Symbolic Verification Gate for LLM Reasoning via Structural Constraint Satisfaction
- CheXPO-v2: Preference Optimization for Chest X-ray VLMs with Knowledge Graph Consistency
- Algorithmic UDAP
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- Collaborative Edge-to-Server Inference for Vision-Language Models
- Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting
- Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering
- Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
- AIAuditTrack: A Framework for AI Security system
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
- Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC
- FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
- dtreg: Describing Data Analysis in Machine-Readable Format in Python and R
- Phythesis: Physics-Guided Evolutionary Scene Synthesis for Energy-Efficient Data Center Design via LLMs
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- LLMs in Interpreting Legal Documents
- Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines
- Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
- CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models
- Calibrated Trust in Dealing with LLM Hallucinations: A Qualitative Study
- Multi-Agent Intelligence for Multidisciplinary Decision-Making in Gastrointestinal Oncology
- Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- Language-driven Fine-grained Retrieval
- Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platforms
- Trusted AI Agents in the Cloud
- LLM Harms: A Taxonomy and Discussion
- The Road of Adaptive AI for Precision in Cybersecurity
- Latent Debate: A Surrogate Framework for Interpreting LLM Thinking
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- The Personalization Paradox: Semantic Loss vs. Reasoning Gains in Agentic AI Q&A
- A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
- Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- Detecting AI Hallucinations in Finance: An Information-Theoretic Method Cuts Hallucination Rate by 92%
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study
- Table as a Modality for Large Language Models
- Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
- Toward a Safe Internet of Agents
- MATCH: Engineering Transparent and Controllable Conversational XAI Systems through Composable Building Blocks
- Factors That Support Grounded Responses in LLM Conversations: A Rapid Review
- REFLEX: Self-Refining Explainable Fact-Checking via Disentangling Truth into Style and Substance
- Enhancing Sequential Recommendation with World Knowledge from Large Language Models
- HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization
- HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
- Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
- What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
- Towards Automating Data Access Permissions in AI Agents
- Synthesizing Precise Protocol Specs from Natural Language for Effective Test Generation
- An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
- Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning
- Build AI Assistants using Large Language Models and Agents to Enhance the Engineering Education of Biomechanics
- Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- Can large language models be a cardinality estimator? An empirical study
- BeautyGuard: Designing a Multi-Agent Roundtable System for Proactive Beauty Tech Compliance through Stakeholder Collaboration
- Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
- Are LLMs The Way Forward? A Case Study on LLM-Guided Reinforcement Learning for Decentralized Autonomous Driving
- On the Influence of Artificial Intelligence on Human Problem-Solving: Empirical Insights for the Third Wave in a Multinational Longitudinal Pilot Study
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- EyeAgent: An Agentic AI System for Multimodal Clinical Decision Support in Ophthalmology
- Unreliable minds, unreliable machines: dyslexic memory, ChatGPT, and the epistemic disobedience of generative AI
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- Trading Vector Data in Vector Databases
- Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
- A Low-Rank Method for Vision Language Model Hallucination Mitigation in Autonomous Driving
- AIA Forecaster: Technical Report
- Evaluation of retrieval-based QA on QUEST-LOFT
- SymLight: Exploring Interpretable and Deployable Symbolic Policies for Traffic Signal Control
- AI-assisted workflow enables rapid, high-fidelity breast cancer clinical trial eligibility prescreening
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- ORCHID: Orchestrated Retrieval-Augmented Classification with Human-in-the-Loop Intelligent Decision-Making for High-Risk Property
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
- KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering
- LLM-enhanced Air Quality Monitoring Interface via Model Context Protocol
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- Using language models to label clusters of scientific documents
- TreeQA: Enhanced LLM-RAG with logic tree reasoning for reliable and interpretable multi-hop question answering
- Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
- A Dual Large Language Models Architecture with Herald Guided Prompts for Parallel Fine Grained Traffic Signal Control
- Retrieval Augmented Generation-Enhanced Distributed LLM Agents for Generalizable Traffic Signal Control with Emergency Vehicles
- Structurally Valid Log Generation using FSM-GFlowNets
- E-Scores for (In)Correctness Assessment of Generative Model Outputs
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- Position: Biology is the Challenge Physics-Informed ML Needs to Evolve
- AnnoBench: A Benchmark for Visualization Annotation Generation
- Machine learning in biological research: key algorithms, applications, and future directions
- From limitation to possibility: affordances and tactical evolution in user–GenAI interaction
- How to Use Generative AI in Educational Research
- Learning from Wildfire Decision Support: large language model analysis of barriers to fire spread in a census of large wildfires in the United States (2011–2023)
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
- BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Hallucination Benchmark for Speech Foundation Models
- Bowling with ChatGPT: On the Evolving User Interactions with Conversational AI Systems
- How well LLM-based test generation techniques perform with newer LLM versions?
- "I Use ChatGPT to Humanize My Words": Affordances and Risks of ChatGPT to Autistic Users
- Mapping and Comparing Climate Equity Policy Practices Using RAG LLM-Based Semantic Analysis and Recommendation Systems
- Ontology-driven generation of parameters for health technology assessment models: a prompt engineering study
- Large-scale manual curation and harmonization of metadata from metagenomic and cancer genomic repositories: challenges and solutions
- What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
- Hallucination Localization in Video Captioning
- Model-Document Protocol for AI Search
- PICOs-RAG: PICO-supported Query Rewriting for Retrieval-Augmented Generation in Evidence-Based Medicine
- M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
- Estimating the Error of Large Language Models at Pairwise Text Comparison
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- InterpDetect: Interpretable Signals for Detecting Hallucinations in Retrieval-Augmented Generation
- DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical services
- MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs Fuzzing
- SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation
- Neural Diversity Regularizes Hallucinations in Language Models
- HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
- FreeChunker: A Cross-Granularity Chunking Framework
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- Enhancing visual-LLM for construction site safety compliance via prompt engineering and Bi-stage retrieval-augmented generation
- LLM-Augmented Symbolic NLU System for More Reliable Continuous Causal Statement Interpretation
- A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
- Learning Efficient and Generalizable Graph Retriever for Knowledge-Graph Question Answering
- ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
- The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
- Beyond "Hallucinations": A Framework for Stable Human-AI Reasoning
- Automated Extraction of Protocol State Machines from 3GPP Specifications with Domain-Informed Prompts and LLM Ensembles
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Element2Vec: Build Chemical Element Representation from Text for Property Prediction
- Counting Hallucinations in Diffusion Models
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG
- ESI: Epistemic Uncertainty Quantification via Semantic-preserving Intervention for Large Language Models
- Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for Recommendation
- Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
- iCodeReviewer: Improving Secure Code Review with Mixture of Prompts
- Credal Transformer: A Principled Approach for Quantifying and Mitigating Hallucinations in Large Language Models
- Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions
- AgentCaster: Reasoning-Guided Tornado Forecasting
- Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
- VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
- Learning to Reason for Hallucination Span Detection
- From Craft to Constitution: A Governance-First Paradigm for Principled Agent Engineering
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- Proof Strategy Extraction from LLMs for Enhancing Symbolic Provers
- CacheClip: Accelerating RAG with Effective KV Cache Reuse
- Failure-Driven Workflow Refinement
- The trust crisis in artificial intelligence: AI hallucinations and human-AI collaboration
- Artificial intelligence in web accessibility: A systematic mapping study
- Speed Kills: Exploring Confused Deputy Attacks Through Edge AI Accelerators
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Fundamentals of Building Autonomous LLM Agents
- Man-Made Heuristics Are Dead. Long Live Code Generators!
- Revisiting Hallucination Detection with Effective Rank-based Uncertainty
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- Stop DDoS Attacking the Research Community with AI-Generated Survey Papers
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Towards Reliable Retrieval in RAG Systems for Large Legal Datasets
- Artificial Intelligence in Detecting Statistical Errors: Implications for Authors, Reviewers, and Editors
- FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
- AgentDR Dynamic Recommendation with Implicit Item-Item Relations via LLM-based Agents
- Domain-Shift-Aware Conformal Prediction for Large Language Models
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- A novel hallucination classification framework
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
- Large Language Models Hallucination: A Comprehensive Survey
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Evaluating Large Language Models for IUCN Red List Species Information
- Geolog-IA: Conversational System for Academic Theses
- Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
- Energy-Regularized Sequential Model Editing on Hyperspheres
- Data Quality Challenges in Retrieval-Augmented Generation
- Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
- Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- Agentic Exploration of Physics Models
- Hallucination is Inevitable for LLMs with the Open World Assumption
- Neural Message-Passing on Attention Graphs for Hallucination Detection
- Adoption paradox of artificial intelligence in computational pathology: a three-stage maturity model from algorithms to clinical integration
- Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
- CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
- Small Language Models for Curriculum-based Guidance
- What If Moderation Didn't Mean Suppression? A Case for Personalized Content Transformation
- Smoothing-Based Conformal Prediction for Balancing Efficiency and Interpretability
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- A model of errors in transformers
- LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
- Are Hallucinations Bad Estimations?
- Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- Generative AI for FFRDCs
- OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
- Augmented and Programmatically Optimized LLM Prompts Reduce Chemical Hallucinations
- Retrieval-Augmented Generation – Ein Erfahrungsbericht und Leitfaden zum Einsatz von wissensbasierten Chatbots an der HU Berlin (Teil 1)
- Evaluating large language models for accuracy incentivizes hallucinations
- How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
- Reflection and Experimental Rigor Are Our AiMS: A New Metacognitive Framework for Experimental Design
- Identifying and Addressing User-level Security Concerns in Smart Homes Using "Smaller" LLMs
- A Knowledge Graph and a Tripartite Evaluation Framework Make Retrieval-Augmented Generation Scalable and Transparent
- Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
- CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation
- EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol
- Campus AI vs. Commercial AI: How Customizations Shape Trust and Usage of LLM as-a-Service Chatbots
- Toward Efficient Influence Function: Dropout as a Compression Tool
- Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- SynBench: A Benchmark for Differentially Private Text Generation
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- SCORE: A Semantic Evaluation Framework for Generative Document Parsing
- Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs
- SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
- Root Cause Analysis of Radiation Oncology Incidents Using Large Language Models
- KoSEL: Knowledge subgraph enhanced large language model for medical question answering
- MMORE: Massive Multimodal Open RAG & Extraction
- HARP: Hallucination Detection via Reasoning Subspace Projection
- Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking
- Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability
- Large Language Models for Security Operations Centers: A Comprehensive Survey
- Virtual Agent Economies
- Robo-Advisors Beyond Automation: Principles and Roadmap for AI-Driven Financial Planning
- Robot guide with multi-agent control and automatic scenario generation with LLM
- Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- Changing the Paradigm from Dynamic Queries to LLM-generated SQL Queries with Human Intervention
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems
- LightAgent: Production-level Open-source Agentic AI Framework
- Deploying AI for Signal Processing education: Selected challenges and intriguing opportunities
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Trees
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
- Temporal Counterfactual Explanations of Behaviour Tree Decisions
- AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
- Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
- Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
- KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino
- Stack Overflow Is Not Dead Yet: Crowd Answers Still Matter
- Prior Distribution and Model Confidence
- Conversational AI increases political knowledge as effectively as self-directed internet search
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Can LLMs Lie? Investigation beyond Hallucination
- DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
- DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
- An LLM-enabled semantic-centric framework to consume privacy policies
- Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- Exploring and Mitigating Fawning Hallucinations in Large Language Models
- TECP: Token-Entropy Conformal Prediction for LLMs
- BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting
- ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Data Auctions for Retrieval Augmented Generation
- Data-driven Discovery of Digital Twins in Biomedical Research
- Multi-Ontology Integration with Dual-Axis Propagation for Medical Concept Representation
- Addressing accuracy and hallucination of LLMs in Alzheimer's disease research through knowledge graphs
- From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
- Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems
- Real-Time Detection of Hallucinated Entities in Long-Form Generation
- A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants
- ConfTuner: Training Large Language Models to Express Their Confidence Verbally
- Skeptik: A Hybrid Framework for Combating Potential Misinformation in Journalism
- Principled Detection of Hallucinations in Large Language Models via Multiple Testing
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Preguss: It Analyzes, It Specifies, It Verifies
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Driving Style Recognition Like an Expert Using Semantic Privileged Information from Large Language Models
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- Applied causality to infer protein dynamics and kinetics
- Can we Evaluate RAGs with Synthetic Data?
- Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- Never compromise with vulnerabilities: a comprehensive survey on AI governance
- MuaLLM: A Multimodal Large Language Model Agent for Circuit Design Assistance with Hybrid Contextual Retrieval-Augmented Generation
- Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
- Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
- GOFAI meets Generative AI: Development of Expert Systems by means of Large Language Models
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
- Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials
- LinguaFluid: Language Guided Fluid Control via Semantic Rewards in Reinforcement Learning
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction Modelling
- StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback
- Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
- LLM-Prior: A Framework for Knowledge-Driven Prior Elicitation and Aggregation
- ASINT: Learning AS-to-Organization Mapping from Internet Metadata
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
- MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
- L3M+P: Lifelong Planning with Large Language Models
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
- HySemRAG: A Hybrid Semantic Retrieval-Augmented Generation Framework for Automated Literature Synthesis and Methodological Gap Analysis
- DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- Latent Knowledge Scalpel: Precise and Massive Knowledge Editing for Large Language Models
- Audio Prototypical Network For Controllable Music Recommendation
- Accessibility Scout: Personalized Accessibility Scans of Built Environments
- How Far Are AI Scientists from Changing the World?
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
- Bridging AI Innovation and Healthcare Needs: Lessons Learned from Incorporating Modern NLP at The BC Cancer Registry
- IM-Chat: A Multi-agent LLM Framework Integrating Tool-Calling and Diffusion Modeling for Knowledge Transfer in Injection Molding Industry
- Probing Information Distribution in Transformer Architectures through Entropy Analysis
- OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
- Large language models provide unsafe answers to patient-posed medical questions
- A Survey of Multimodal Hallucination Evaluation and Detection
- Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery
- Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
- Explainable Mapper: Charting LLM Embedding Spaces Using Perturbation-Based Explanation and Verification Agents
- Can generative AI detect and fix real-world cryptographic misuses?
- Byzantine-Robust Decentralized Coordination of LLM Agents
- The Promise of Foundational Large Language Models in Analysis and Interpretation of Wearable Data: Implications for Physical Behavior Research
- InsightX Agent: An LMM-based Agentic Framework with Integrated Tools for Reliable X-ray NDT Analysis
- Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
- Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization
- Theoretical Foundations and Mitigation of Hallucination in Large Language Models
- Can generative AI reliably synthesise literature? exploring hallucination issues in ChatGPT
- AC4A: Access Control for Agents
- Designing Conversational AI to Support Think-Aloud Practice in Technical Interview Preparation for CS Students
- ElectriQ: A Benchmark for Assessing the Response Capability of Large Language Models in Power Marketing
- Large Language Models in the Travel Domain: An Industrial Experience
- Impact of Code Context and Prompting Strategies on Automated Unit Test Generation with Modern General-Purpose Large Language Models
- Critique of impure reason: Unveiling the reasoning behaviour of medical large language models
- TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment
- BifrostRAG: Bridging Dual Knowledge Graphs for Multi-Hop Question Answering in Construction Safety
- Hyperbolic Deep Learning for Foundation Models: A Survey
- From Extraction to Synthesis: Entangled Heuristics for Agent-Augmented Strategic Reasoning
- Symbiotic Agents: A Novel Paradigm for Trustworthy AGI-driven Networks
- Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs
- Probabilistic distances-based hallucination detection in LLMs with RAG
- PALM: Synergizing Program Analysis and LLMs to Enhance Rust Unit Test Coverage
- EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements
- FedRAG: A Framework for Fine-Tuning Retrieval-Augmented Generation Systems
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- AnalogTester: A Large Language Model-Based Framework for Automatic Testbench Generation in Analog Circuit Design
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
- WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis
- How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs
- Multi-Agent Retrieval-Augmented Framework for Evidence-Based Counterspeech Against Health Misinformation
- Evaluating Retrieval-Augmented Generation Agents for Autonomous Scientific Discovery in Astrophysics
- A Mathematical Theory of Discursive Networks
- Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
- SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads
- CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs
- Real-Time AI-Driven Pipeline for Automated Medical Study Content Generation in Low-Resource Settings: A Kenyan Case Study
- LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers
- Data Discovery using LLMs -- A Study of Data User Behaviour
- KEA Explain: Explanations of Hallucinations using Graph Kernel Analysis
- A Concept for Autonomous Problem-Solving in Intralogistics Scenarios
- ReservoirChat: Interactive Documentation Enhanced with LLM and Knowledge Graph for ReservoirPy
- STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
- Escaping Mode Collapse in LLM Generation via Geometric Regulation
- ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Computer Science Education in the Age of Generative AI
- Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
- Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
- Shiksha Copilot: Teacher-AI Collaboration for Curating and Customizing Lesson Plans in Low-Resource Schools
- A Clinically-Grounded Two-Stage Framework for Renal CT Report Generation
- Multimodal Large Language Model With Knowledge Retrieval Using Flowchart Embedding for Forming Follow-Up Recommendations for Pancreatic Cystic Lesions
- Do LLMs Dream of Discrete Algorithms?
- Safe and Performant Deployment of Autonomous Systems via Model Predictive Control and Hamilton-Jacobi Reachability Analysis
- VALID-Mol: a Systematic Framework for Validated LLM-Assisted Molecular Design
- From Data Center IoT Telemetry to Data Analytics Chatbots -- Virtual Knowledge Graph is All You Need
- A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis
- Adapting University Policies for Generative AI: Opportunities, Challenges, and Policy Solutions in Higher Education
- Weak-to-Strong GraphRAG: Aligning Weak Retrievers with Large Language Models for Graph-based Retrieval Augmented Generation
- KScope: A Framework for Characterizing the Knowledge Status of Language Models
- EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora
- Exploring the Effects of Chatbot Anthropomorphism and Human Empathy on Human Prosocial Behavior Toward Chatbots
- OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
- PromptAug: Fine-grained Conflict Classification Using Data Augmentation
- KnowML: Improving Generalization of ML-NIDS with Attack Knowledge Graphs
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- Hallucination Detection with Small Language Models
- Prover Agent: An Agent-Based Framework for Formal Mathematical Proofs
- Dialogic Pedagogy for Large Language Models: Aligning Conversational AI with Proven Theories of Learning
- AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
- Advanced Applications of Generative AI in Actuarial Science: Case Studies Beyond ChatGPT
- Refine Medical Diagnosis Using Generation Augmented Retrieval and Clinical Practice Guidelines
- CARTS: Collaborative Agents for Recommendation Textual Summarization
- HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models
- UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
- SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2
- PCS Workflow for Veridical Data Science in the Age of AI
- Research on Graph-Retrieval Augmented Generation Based on Historical Text Knowledge Graphs
- StorySage: Conversational Autobiography Writing Powered by a Multi-Agent Framework
- InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
- The Structure-Content Trade-off in Knowledge Graph Retrieval
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- QuantMCP: Grounding Large Language Models in Verifiable Financial Reality
- Designing GenAI Tools for Personalized Learning Implementation: Theoretical Analysis and Prototype of a Multi-Agent System
- Augmenting the Generality and Performance of Large Language Models for Software Engineering
- A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis
- Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
- Revealing the True Indicators: Understanding and Improving IoC Extraction From Threat Reports
- LLM Embedding-based Attribution (LEA): Quantifying Source Contributions to Generative Model's Response for Vulnerability Analysis
- Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic Retrieval-Augmented Generation for Industry Challenges
- Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge
- Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
- Explainability in Context: A Multilevel Framework Aligning AI Explanations with Stakeholder with LLMs
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions
- Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation
- Large Language Models are Good Relational Learners
- Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
- From First Draft to Final Insight: A Multi-Agent Approach for Feedback Generation
- Benchmarking Large Language Models on Homework Assessment in Circuit Analysis
- One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
- Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models
- On the Fundamental Impossibility of Hallucination Control in Large Language Models
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- A Large Language Model for Feasible and Diverse Population Synthesis
- CDE-Mapper: Using Retrieval-Augmented Language Models for Linking Clinical Data Elements to Controlled Vocabularies
- Graph Drawing for LLMs: An Empirical Evaluation
- ScoreRAG: A Retrieval-Augmented Generation Framework with Consistency-Relevance Scoring and Structured Summarization for News Generation
- AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
- Hallucination to Consensus: Multi-Agent LLMs for End-to-End JUnit Test Generation
- Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection
- V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving
- An Iterative Question-Guided Framework for Knowledge Base Question Answering
- MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
- LAMARL: LLM-Aided Multi-Agent Reinforcement Learning for Cooperative Policy Generation
- KG-TRACES: Enhancing Large Language Models with Knowledge Graph-constrained Trajectory Reasoning and Attribution Supervision
- A Hashgraph-Inspired Consensus Mechanism for Reliable Multi-Model Reasoning
- A Large Language Model Based Pipeline for Review of Systems Entity Recognition from Clinical Notes
- On Early Detection of Hallucinations in Factual Question Answering
- Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning
- Fewer Hallucinations, More Verification: A Three-Stage LLM-Based Framework for ASR Error Correction
- Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
- Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model
- Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding
- From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
- Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
- AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive Reasoning
- Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation
- If Pigs Could Fly... Can LLMs Logically Reason Through Counterfactuals?
- Read Your Own Mind: Reasoning Helps Surface Self-Confidence Signals in LLMs
- D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
- CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models
- Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
- RedAHD: Reduction-Based End-to-End Automatic Heuristic Design with Large Language Models
- DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems
- Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
- InFact: Informativeness Alignment for Improved LLM Factuality
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- Direct Retrieval-augmented Optimization: Synergizing Knowledge Selection and Language Models
- Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention Heads
- Architectures of Error: A Philosophical Inquiry into AI and Human Code Generation
- LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
- Do Large Language Models (Really) Need Statistical Foundations?
- Removal of Hallucination on Hallucination: Debate-Augmented RAG
- Formally Solving Answer-Construction Problems in Lean
- Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are fully accurate
- Seek-CAD: A Self-refined Generative Modeling for 3D Parametric CAD Using Local Inference via DeepSeek
- EVADE: Multimodal Benchmark for Evasive Content Detection in E-Commerce Applications
- Retrieval Augmented Generation-based Large Language Models for Bridging Transportation Cybersecurity Legal Knowledge Gaps
- The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
- Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
- Personalizing Student-Agent Interactions Using Log-Contextualized Retrieval Augmented Generation (RAG)
- Cracking Aegis: An Adversarial LLM-based Game for Raising Awareness of Vulnerabilities in Privacy Protection
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
- Rethinking Code Review Workflows with LLM Assistance: An Empirical Study
- Automated Feedback Loops to Protect Text Simplification with Generative AI from Information Loss
- O2-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
- Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs
- LLM-Powered AI Agent Systems and Their Applications in Industry
- A Survey on the Application of Large Language Models in Scenario-Based Testing of Automated Driving Systems
- Integral Imprecise Probability Metrics
- Conformal Language Model Reasoning with Coherent Factuality
- UniErase: Towards Balanced and Precise Unlearning in Language Models
- InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
- SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
- Know When to Abstain: Optimal Selective Classification with Likelihood Ratios
- Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
- BugRepro: Enhancing Android Bug Reproduction with Domain-Specific Knowledge Integration
- Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
- Concept Incongruence: An Exploration of Time and Death in Role Playing
- Enhancing LLMs via High-Knowledge Data Selection
- MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
- Reasoning Models Better Express Their Confidence
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning
- LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries
- Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review
- Incentivizing Truthful Language Models via Peer Elicitation Games
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
- Know Or Not: a library for evaluating out-of-knowledge base robustness
- Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments
- Truth Neurons
- BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs
- Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
- Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
- Fine-grained Contrastive Learning for ECG-Report Alignment with Waveform Enhancement
- Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
- EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
- Phare: A Safety Probe for Large Language Models
- From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness
- Terminators: Terms of Service Parsing and Auditing Agents
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV Cache
- Eliminating Hallucination-Induced Errors in LLM Code Generation with Functional Clustering
- Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
- Safety Invariants for Agents Orchestrating Irreversible State Transitions: A Four-Dimensional Formalism Evaluated on Public Ledgers
- Campus AI vs Commercial AI: A Late-Breaking Study on How LLM As-A-Service Customizations Shape Trust and Usage Patterns
- StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
- AI-enhanced semantic feature norms for 786 concepts
- Atomic Consistency Preference Optimization for Long-Form Question Answering
- An AI-Powered Research Assistant in the Lab: A Practical Guide for Text Analysis Through Iterative Collaboration with LLMs
- Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generation
- Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies
- Securing RAG: A Risk Assessment and Mitigation Framework
- CellTypeAgent: Trustworthy cell type annotation with Large Language Models
- Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency
- Vision Foundation Model Embedding-Based Semantic Anomaly Detection
- Multimodal Cancer Modeling in the Age of Foundation Model Embeddings
- Examining the Role of LLM-Driven Interactions on Attention and Cognitive Engagement in Virtual Classrooms
- RAI: Flexible Agent Framework for Embodied AI
- Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent
- POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
- LLM-Flock: Decentralized Multi-Robot Flocking via Large Language Models and Influence-Based Consensus
- Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
- Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification
- Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
- Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community
- ”My AI is Lying to Me”: User-reported LLM hallucinations in AI mobile apps reviews
- Prompt engineering for bibliographic web-scraping
- SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
- Interpretable graph-based models on multimodal biomedical data integration: A technical review and benchmarking
- ReLI: A Language-Agnostic Approach to Human-Robot Interaction
- Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM
- CaGR-RAG: Context-aware Query Grouping for Disk-based Vector Search in RAG Systems
- To Trust or Not to Trust: Authors' Response to AI-based Reviews
- LLM Agents in Law: Taxonomy, Applications, and Challenges
- The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale
- Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
- GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
- Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
- Red Teaming Large Language Models for Healthcare
- Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems
- Toward Epistemic Stability: Engineering Consistent Procedures for Industrial LLM Hallucination Reduction
- Stan: An LLM-based thermodynamics course assistant
- Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity
- Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting
- Grid-Mind: An LLM-Orchestrated Multi-Fidelity Agent for Automated Connection Impact Assessment
- Spilled Energy in Large Language Models
- A New Strategy for Artificial Intelligence: Training Foundation Models Directly on Human Brain Data
- Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
- Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
- Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
- Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
- Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction
- Adjust to reality: LLM-driven test-time semantic adjustment for zero-shot fault diagnosis
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
- HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation
- kAgent: An execution-guided crash resolution agent for the Linux kernel
- If Concept Bottlenecks are the Question, are Foundation Models the Answer?
- Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator
- Hallucination in World Models is Predictable and Preventable
- PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
- TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
- Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
- Scalable Token-Level Hallucination Detection in Large Language Models
- Reversible Lifelong Model Editing via Semantic Routing-Based LoRA
- An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
- The Algorithmic Self-Portrait: Deconstructing Memory in ChatGPT
- MIND: Empowering Mental Health Clinicians with Multimodal Data Insights through a Narrative Dashboard
- Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
- Detect, Explain, Escalate: Sustainable Dialogue Breakdown Management for LLM Agents
- Toward Safe and Human-Aligned Game Conversational Recommendation via Multi-Agent Decomposition
- A Conditional Companion: Lived Experiences of People with Mental Health Disorders Using LLMs
- VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
- aiPlato: A Novel AI Tutoring and Step-wise Feedback System for Physics Homework
- ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction
- Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores
- ROAD: Reflective Optimization via Automated Debugging for Zero-Shot Agent Alignment
- Exploring General-Purpose Autonomous Multimodal Agents for Pathology Report Generation
- AI Awareness
- HalluLens: LLM Hallucination Benchmark
- Transforming remanufacturing automation with large language models: A forward-looking analysis with case studies
- Can ChatGPT Learn My Life From a Week of First-Person Video?
- Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
- Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
- ArXivBench: When You Should Avoid Using ChatGPT for Academic Writing
- Exploring the Role of Large Language Models in Cybersecurity: A Systematic Survey
- FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation
- aiXamine: Simplified LLM Safety and Security
- Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
- PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
- Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
- ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
- Toward Generation of Test Cases from Task Descriptions via History-aware Planning
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
- A Comprehensive Survey on LLM‐Based Network Management and Operations
- Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
- To trust or not to trust a human(-like) AI—A scoping review and conjoint analyses on factors influencing anthropomorphism and trust
- Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
- CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
- Enhancing Autonomous Driving Systems with On-Board Deployed Large Language Models
- Hallucination-Aware Generative Pretrained Transformer for Cooperative Aerial Mobility Control
- DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
- A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science
- DioR: Adaptive Cognitive Detection and Contextual Retrieval Optimization for Dynamic Retrieval-Augmented Generation
- Can LLMs Assist Expert Elicitation for Probabilistic Causal Modeling?
- C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
- How new data permeates LLM knowledge and how to dilute it
- HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
- MedHal: An Evaluation Dataset for Medical Hallucination Detection
- Out of Style: RAG's Fragility to Linguistic Variation
- Robust Hallucination Detection in LLMs via Adaptive Token Selection
- Enhancing Large Language Models through Neuro-Symbolic Integration and Ontological Reasoning
- Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
- Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
- Large language models as uncertainty-calibrated optimizers for experimental discovery
- Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
- GraphRAFT: Retrieval Augmented Fine-Tuning for Knowledge Graphs on Graph Databases
- Beyond Answers: How LLMs Can Pursue Strategic Thinking in Education
- Autono: A ReAct-Based Highly Robust Autonomous Agent Framework
- Generative AI Enhanced Financial Risk Management Information Retrieval
- AD-GPT: Large Language Models in Alzheimer's Disease
- Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
- CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
Discussions
- Quellen: dl.acm.org/doi/full/10...., arxiv.org/html/2509.22..., help.openai.com/en/articles/... [bsky, 17 points, 3 comments]
- Spannendes Uni-Projekt. KI-Hallizunationen: Ursachen, Wirkung, Lösungen. Ich bin Team "Ursachen". Das Paper schlägt eine strukturierte Taxonomie vor: Es gliedert Halluzinationen in Kategorien, um Klar [bsky, 9 points, 1 comments]
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions https://arxiv.org/abs/2311.05232v1 #AI #hallucination [bsky, 1 points, 0 comments]
- arxiv.org/abs/2311.05232 [bsky, 0 points, 1 comments]
Related