Hallucination is Inevitable: An Innate Limitation of Large Language Models
2024/01/22 by Ziwei Xu, Sanjay Jain, Xu, Ziwei +3 · 42 voices · 94 citations
Computer Science · Engineering · #Ferroelectric and Negative Capacitance Devices #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2401.11817
openalex publication_date 2024/01/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Hallucination has been widely recognized to be a significant drawback for large language models (LLMs). There have been many works that attempt to reduce the extent of hallucination. These efforts have mostly been empirical so far, which cannot answer the fundamental question whether it can be completely eliminated. In this paper, we formalize the problem and show that it is impossible to eliminate hallucination in LLMs. Specifically, we define a formal world where hallucination is defined as inconsistencies between a computable LLM and a computable ground truth function. By employing results from learning theory, we show that LLMs cannot learn all the computable functions and will therefore inevitably hallucinate if used as general problem solvers. Since the formal world is a part of the real world which is much more complicated, hallucinations are also inevitable for real world LLMs. Furthermore, for real world LLMs constrained by provable time complexity, we describe the hallucination-prone tasks and empirically validate our claims. Finally, using the formal world framework, we discuss the possible mechanisms and efficacies of existing hallucination mitigators as well as the practical implications on the safe deployment of LLMs.
Cited by
- Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Information-Theoretic Limits of Reliability and Scaling in Language Models
- Language Model Teams as Distributed Systems
- Epistemological Fault Lines Between Human and Artificial Intelligence
- AI, Digital Platforms, and the New Systemic Risk
- Hallucination Stations: On Some Basic Limitations of Transformer-Based Language Models
- Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
- TARGET: Automated Scenario Generation from Traffic Rules for Testing Autonomous Vehicles via Validated LLM-Guided Knowledge Extraction
- A Taxonomy of Confabulations and the Perception-Reality Gap in LLM-Assisted Immersive Scene Editing
- Specula: Scaling formal specifications for autonomous model checking of system code
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- COMPARE: Clinical Optimization with Modular Planning and Assessment via RAG-Enhanced AI-OCT: Superior Decision Support for Percutaneous Coronary Intervention Compared to ChatGPT-5 and Junior Operators
- Calibrated Trust in Dealing with LLM Hallucinations: A Qualitative Study
- Insured Agents: A Decentralized Trust Insurance Mechanism for Agentic Economy
- Latent Debate: A Surrogate Framework for Interpreting LLM Thinking
- DialogGuard: Multi-Agent Psychosocial Safety Evaluation of Sensitive LLM Responses
- Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming
- Build AI Assistants using Large Language Models and Agents to Enhance the Engineering Education of Biomechanics
- Cost-Driven Synthesis of Sound Abstract Interpreters
- Reducing Hallucinations in LLM-Generated Code via Semantic Triangulation
- A Neurosymbolic Approach to Natural Language Formalization and Verification
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- Stemming Hallucination in Language Models Using a Licensing Oracle
- AI as We Describe It: How Large Language Models and Their Applications in Health are Represented Across Channels of Public Discourse
- Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
- A Criminology of Machines
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
- ATLAS: Harnessing retrieval-augmented generation (RAG)
- Group size effects and collective misalignment in LLM multi-agent systems
- PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
- Neural Diversity Regularizes Hallucinations in Small Models
- Interpretable Question Answering with Knowledge Graphs
- The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
- Automated Extraction of Protocol State Machines from 3GPP Specifications with Domain-Informed Prompts and LLM Ensembles
- Towards Human-Centric Intelligent Treatment Planning for Radiation Therapy
- RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
- Large Language Models Hallucination: A Comprehensive Survey
- Increasing LLM response trustworthiness using voting ensembles
- Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
- Improving Metacognition and Uncertainty Communication in Language Models
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- Black-box Context-free Grammar Inference for Readable & Natural Grammars
- Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
- Hallucination is Inevitable for LLMs with the Open World Assumption
- Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
- Bridging Language Models and Formal Methods for Intent-Driven Optical Network Design
- DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
- Limitations on Accurate, Trusted, Human-level Reasoning
- Are Hallucinations Bad Estimations?
- An LLM-based Agentic Framework for Accessible Network Control
- Asymmetric Communication: Large Language Models and Language Games
- Campus AI vs. Commercial AI: How Customizations Shape Trust and Usage of LLM as-a-Service Chatbots
- SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
- DPCheatSheet: Using Worked and Erroneous LLM-usage Examples to Scaffold Differential Privacy Implementation
- Shapes of Cognition for Computational Cognitive Modeling
- Regulating genome language models: navigating policy challenges at the intersection of AI and genetics
- Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability
- Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- Temporal Counterfactual Explanations of Behaviour Tree Decisions
- Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification
- Mitigating Harmful Erraticism in LLMs Through Dialectical Behavior Therapy Based De-Escalation Strategies
- What Would an LLM Do? Evaluating Policymaking Capabilities of Large Language Models
- Can LLMs Lie? Investigation beyond Hallucination
- Artificial or Human Intelligence?
- Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning
- CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
- Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization
- Assessing student perceptions and use of instructor versus <scp>AI</scp> ‐generated feedback
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- Hallucinations in medical devices
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
- Considering the ethics of large machine learning models in the chemical sciences
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- The Problem of Atypicality in LLM-Powered Psychiatry
- Dean of LLM Tutors: Exploring Comprehensive and Automated Evaluation of LLM-generated Educational Feedback via LLM Feedback Evaluators
- Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
- FAIRJupyter4AI: A Corpus of Computational Notebooks for AI
- A Survey on Data Security in Large Language Models
- T-GRAG: A Dynamic GraphRAG Framework for Resolving Temporal Conflicts and Redundancy in Knowledge Retrieval
- Latent Knowledge Scalpel: Precise and Massive Knowledge Editing for Large Language Models
- Visual Language Models as Zero-Shot Deepfake Detectors
- OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
- Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
- NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
Discussions
- Hallucination is inevitable: An innate limitation of large language models [hn, 308 points, 474 comments]
- "In this paper, we formalize the problem and show that it is impossible to eliminate hallucination in LLMs." Hallucination is Inevitable: An Innate Limitation of Large Language Models arxiv.org/pdf/24 [bsky, 99 points, 4 comments]
- Hallucinations are inevitable and likely an unsolvable problem. "By employing results from learning theory, we show that LLMs cannot learn all the computable functions and will therefore inevitably [bsky, 68 points, 1 comments]
- aber wir sind sooooooooooo 🤏 kurz davor dass die klüger als menschen sind [bsky, 35 points, 3 comments]
- Hallucination is Inevitable: An Innate Limitation of Large Language Models [lobsters, 24 points, 17 comments]
- Hallucination Is Inevitable: An Innate Limitation of Large Language Models (2025) [hn, 14 points, 11 comments]
- This is one of them: arxiv.org/abs/2401.11817 Mind you: a proof that there are classes of problems where token generation from a model cannot be correct, not a proof for all hallucination causes. [bsky, 9 points, 0 comments]
- This one as well. The math isn't that hard either. The point being that hallucinations are a structurally inevitable outcome of LLMs. They run up against Gödel's incompleteness theorem & the Halting P [bsky, 9 points, 2 comments]
- arxiv.org/abs/2401.11817 [bsky, 6 points, 1 comments]
- and also, there seems to be literally a formal proof LLM will always hallucinate: arxiv.org/abs/2401.11817 [bsky, 4 points, 1 comments]
- arxiv.org/abs/2401.11817 Where's your citation for your claim? [bsky, 4 points, 1 comments]
- Researchers have been saying this since at least the beginning of 2024, but I guess it's news when OpenAI also says it? arxiv.org/abs/2401.11817 [bsky, 3 points, 1 comments]
- Hallucination Is Inevitable: An Innate Limitation of Large Language Models [hn, 3 points, 2 comments]
- Here's one paper I've seen on it arxiv.org/abs/2401.11817 But also just given the way LLMs function, they're designed to reproduce patterns from past data. They have no knowledge of the present contex [bsky, 3 points, 0 comments]
- Studies show that today's #AI based on large language models can't avoid "hallucinating" — or what we might better call "miraging." See Xu et al doi.org/10.48550/arX... and @irisvanrooij.bsky.social e [bsky, 3 points, 0 comments]
- Hallucination is inevitable: An innate limitation of large language models (arxiv.org) Main Link | Discussion [bsky, 2 points, 0 comments]
- Shitposting might reduce output quality, but hallucinations are proven to be baked into any LLM, regardless of data quality arxiv.org/abs/2401.11817 [bsky, 2 points, 0 comments]
- "All LLMs will hallucinate...Without guardrail & fences, LLMs cannot be used for critical decision making. This is a direct corollary of the conclusion above...Without human control, LLMs cannot be us [bsky, 2 points, 1 comments]
- I think the main problem is that hallucinations are often part of a broader story. For very concrete questions I'd think you could do this, but I imagine that finding incongruous elements in larger st [bsky, 2 points, 0 comments]
- not a huge fan of mocking people while they are coping but I have seen papers like this in Real Life e.g. arxiv.org/abs/2401.11817 [bsky, 2 points, 2 comments]
- There's a good paper on why LLMs error on these problems and by design always will. Anything with a large permutation space is too expensive to represent and gets pruned. arxiv.org/pdf/2401.11817 Pret [bsky, 1 points, 1 comments]
- OpenAI no admite nada. Como casi siempre, son los investigadores los que se preocupan de estas cosas. arxiv.org/abs/2401.11817 Y el artículo ya lleva tiempo en revisión, no es de ahora. [bsky, 1 points, 2 comments]
- Apparently it’s mathematically impossible to eliminate hallucinations from AI arxiv.org/abs/2401.11817 [bsky, 1 points, 0 comments]
- Hallucinations are an inherent property of LLMs. There are Youtube videos by Yann LeCun discussing this (one of the pre-eminent AI researchers), but there's also this paper if you prefer text: arxiv.o [bsky, 1 points, 1 comments]
- #IA Les hallucinations sont inévitables: une limitation innée des grands modèles de langages arxiv.org/abs/2401.11817 [bsky, 1 points, 0 comments]
- Unlike an ai I an aware of the bounds of my knowledge and happily defer to those more invested in a field than I am. arxiv.org/abs/2401.11817 [bsky, 0 points, 1 comments]
- Hallucination is Inevitable: An Innate Limitation of Large Language Models arxiv.org/pdf/2401.118... [bsky, 0 points, 0 comments]
- L'intervention humaine n'est pas un choix « pragmatique » pour ce qui est des LLM, puisque les hallucinations y sont absolument inévitables avant peaufinage manuel : arxiv.org/abs/2401.11817 Ce qui n' [bsky, 0 points, 0 comments]
- Hallucination is inevitable: An innate limitation of large language models arxiv.org/abs/2401.11817 Resistance is futile; #AI #LLM [bsky, 0 points, 0 comments]
- @rstockm.bsky.social Kennst du das paper hier schon: arxiv.org/pdf/2401.118... ? Die Behaptung ist, dass alle LLMs halluzinieren müssen. Darin ist auch ein Beispiel wie man prompts generieren kann an [bsky, 0 points, 0 comments]
- Make shit up 25% of the time? The rate of error is unknown but LLMs can't not make shit up its fundamental to their design they are not appropriate tools for a medical setting that's without even cons [bsky, 0 points, 0 comments]
- That guy's still at it? arxiv.org/abs/2401.11817 [bsky, 0 points, 0 comments]
- The famous Chinchilla paper shows doubling model size also requires doubling training tokens, so compute does represent a constraint: arxiv.org/abs/2203.15556. But more compute alone won’t fix it, we [bsky, 0 points, 1 comments]
- Nee, het is veel fundamenteler. Het gaat om hoe die modellen technisch werken. Hallucinaties zijn feitelijk onvermijdelijk. arxiv.org/abs/2401.11817 [bsky, 0 points, 1 comments]
- “specialists see it as a threat” sounds like the NYT talking about RFK’s vaccine attitudes lol TBC whether the confidently wrong point is a “current design flaw” or the basis of the tech arxiv.org/abs [bsky, 0 points, 1 comments]
- Do we really have to point you to every published peer-reviewed paper demonstrating that LLMs will *always* hallucinate convincing but fundamentally incorrect 'answers', and this is not a solvable pro [bsky, 0 points, 0 comments]
- Hallucination Is Inevitable: An Innate Limitation of Large Language Models https://arxiv.org/abs/2401.11817 [bsky, 0 points, 0 comments]
- via Mike Tyka: Hallucination is Inevitable: An Innate Limitation of Large Language Models https://arxiv.org/abs/2401.11817 "In this paper, we formalize the problem and show that it is impossible to [bsky, 0 points, 1 comments]
- Well then you should publish a paper disputing this one! arxiv.org/abs/2401.11817 [bsky, 0 points, 1 comments]
- Hallucination Is Inevitable: An Innate Limitation of Large Language Models https:// arxiv.org/abs/2401.11817 # arxiv [mastodon, 0 points, 0 comments]
- My language attempts to correct for this. Coincidental it’s also called Innate. HTTPS://github.com/n8k99/innatescript ↩ thought-police [bsky, 0 points, 0 comments]
- OTOH, arxiv.org/pdf/2401.118... [bsky, 0 points, 1 comments]
Related