Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
2025/09/04 by Adam Tauman Kalai, Ofir Nachum, Kalai, Adam Tauman +5 · 78 voices · 76 citations
Computer Science · #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2509.04664
Abstract
Layer-0 “suppressor” heads explain why LMs trade factuality for hedging. In GPT-2 Medium, ablating heads 0:2, 0:4, 0:7 increases logit-difference (ΔLD) by 0.40–0.85 across four single-token probes and improves calibration (ECE 0.122 → 0.091). Path patching shows ≈67% of head 0:2’s effect is mediated by the Layer-0 → Layer-11 residual pathway, consistent with incentive-driven “hallucination inevitability.” Mistral-7B exhibits an architecture-adapted variant. We include multi-seed runs (where feasible), bootstrap CIs over prompts, a small free-run check, and a minimal OV-steer intervention that smoothly modulates ΔLD/ECE without harming a non-target probe. Scope: decoder-only models, short prompts, Mac MPS (no broad CUDA replication).
Citations
Cited by
- Enhancing SLMs for Sustainable Code Optimization in Radio-Astronomy
- Memoir: Should a Model Write to Its Memory While It Thinks?
- Mechanistic Attention Guidance for Agent Memory Refinement
- Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
- The CRAFT principles for the responsible use of large language models in policymaking
- Inductive Risk of AI Hype
- Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering
- Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
- Training LLMs for Honesty via Confessions
- LLMs can hide text in other text of the same length
- Smart Microscopy: Current Implementations and a Roadmap for Interoperability
- Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
- From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
- Cognitive Dark Matter: Measuring What AI Misses
- Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- Knowing Your Uncertainty -- On the application of LLM in social sciences
- Hallucinated citations produced by generative artificial intelligence may constitute research misconduct when citations function as data in scholarly papers
- Latent Debate: A Surrogate Framework for Interpreting LLM Thinking
- TaskEval: Synthesised Evaluation for Foundation-Model Tasks
- Small Language Models Reshape Higher Education: Courses, Textbooks, and Teaching
- Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
- Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning
- Plan-X: Instruct Video Generation via Semantic Planning
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Toward Honest Language Models for Deductive Reasoning
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
- Unreliable minds, unreliable machines: dyslexic memory, ChatGPT, and the epistemic disobedience of generative AI
- Making LLMs Reliable When It Matters Most: A Five-Layer Architecture for High-Stakes Decisions
- When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
- Language Generation: Complexity Barriers and Implications for Learning
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
- SoraNav: Adaptive UAV Task-Centric Navigation via Zeroshot VLM Reasoning
- ATA: A Neuro-Symbolic Approach to Implement Autonomous and Trustworthy Agents
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
- PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- An Expert-grounded benchmark of General Purpose LLMs in LCA
- Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
- JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
- Beyond "Hallucinations": A Framework for Stable Human-AI Reasoning
- The Transformation of Epistemic Agency and Governance in Higher Education Through Large Language Models: Toward a future of organized immaturity
- RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge
- Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
- Fast Leave-One-Out Approximation from Fragment-Target Prevalence Vectors (molFTP) : From Dummy Masking to Key-LOO for Leakage-Free Feature Construction
- InvThink: Premortem Reasoning for Safer Language Models
- Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
- Improving Metacognition and Uncertainty Communication in Language Models
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Hallucination is Inevitable for LLMs with the Open World Assumption
- Automated Vulnerability Validation and Verification: A Large Language Model Approach
- KI bei Studierenden: Empirische Erhebung zu Nutzungspraxis, Erwartungen und Haltungen
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Are Hallucinations Bad Estimations?
- Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
- ALIMA – Ein RAG-basiertes System zur LLM-gestützten Sacherschließung: Prototypentwicklung und erste Erfahrungen aus der Praxis
- Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Disproving the Feasibility of Learned Confidence Calibration Under Binary Supervision: An Information-Theoretic Impossibility
- Evaluating Large Language Models for Evidence-Based Clinical Question Answering
- XML Prompting as Grammar-Constrained Interaction: Fixed-Point Semantics, Convergence Guarantees, and Human-AI Protocols
- Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification
- PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
- Grounding the Ungrounded: A Spectral-Graph Framework for Quantifying Hallucinations in Multimodal LLMs
Discussions
- The OpenAi preprint on arXiv arxiv.org/pdf/2509.04664 [bsky, 134 points, 1 comments]
- Published paper proving that #ChatGPT will always make things up. Not sometimes. Not until the next update. Always. They proved it with math. Even with perfect data and unlimited computing power, AI m [bsky, 94 points, 4 comments]
- Arxiv link to the paper referenced in the article arxiv.org/pdf/2509.04664 [bsky, 61 points, 1 comments]
- While the C suite at OpenAI would like you to believe that they are about to break new ground on intelligent agents, their researchers tell a very different story. This paper explains that AI hallucin [bsky, 49 points, 1 comments]
- Users would leave overnight. So a fix exists, but it would kill the product. arxiv.org/abs/2509.04664 [bsky, 18 points, 2 comments]
- What looks to be an interesting paper (outside of my area of expertise) on LLMs and hallucinations: "hallucinations persist due to the way most evaluations are graded -- language models are optimized [bsky, 14 points, 2 comments]
- Το Cornell έβγαλε το ακόλουθο paper στο οποιο καταλήγει οτι οσο εξυπνότερο ειναι ενα ΑΙ μοντελο τοσο μεγαλυτερα ψεματα θα πει.Και αυτο δεν εχει ταβανι γιατι ενω τα "πιο χαζα" μοντελα δεν απαντουσαν αν [bsky, 11 points, 1 comments]
- I guess it's kind of obvious, given the design, that LLMs will hallucinate, but interesting to see Open AI researchers showing that the "majority of mainstream evaluations reward hallucinatory behavio [bsky, 9 points, 1 comments]
- yes, it’s pretty well-known. I think the actual paper has a good point though (arxiv.org/pdf/2509.04664), namely that until you fix the benchmarks it’s not gonna get better [bsky, 8 points, 1 comments]
- arxiv.org/abs/2509.04664 [bsky, 7 points, 2 comments]
- Your LLM Won’t Stop Lying Any Time Soon [lemmy, 6 points, 2 comments]
- they will hallucinate even if trained on perfect data arxiv.org/pdf/2509.04664 [bsky, 6 points, 1 comments]
- Uusi tämä tieto ei ole, mutta onhan se hyvä, että maailman suurimman kielimalliyrityksenkin tutkijat tulevat tähän lopputulokseen. Tutkijoiden mukaan valuvika pysyy kielimalleissa kehittäisi niitä kui [bsky, 5 points, 1 comments]
- Original study 🧪🤖 [bsky, 5 points, 0 comments]
- It’s a function of incentive programming OpenAI researchers themselves recently concluded you could incentivise a LLM to admit it doesn’t know, but that would limit its binary determinations to set pr [bsky, 4 points, 1 comments]
- Anyone still reporting on AI without citing how both OpenAI & Apple dropped papers debunking the industry's claims is derelict in their journalistic duties. That info should be in the SEO footer they [bsky, 4 points, 1 comments]
- LLMs don’t lie on purpose—they guess when unsure because today’s benchmarks reward confident answers, not admitting uncertainty. Hallucinations arise from this misalignment. To fix AI trustworthiness, [bsky, 4 points, 0 comments]
- ChatGPT’s engineers have even published a paper explaining why eliminating hallucinations is mathematically impossible. There is a lot we don’t know about how our brains work, but we do know they don’ [bsky, 4 points, 1 comments]
- Hehehe you mean you aren't going to take openAI at their word that hallucinations "aren't a problem with the model man *takes a big drag on something*... hallucinations are a problem with HUMANS judgi [bsky, 3 points, 2 comments]
- Q: What I find interesting about #AI research is that it can’t help but fall back on human concepts (eg ‘hallucination’) to explain how things (don’t) work. Yet authors insist that hallucinations for [bsky, 3 points, 1 comments]
- Straight from the people making AI: All AI systems make stuff up because they were (inadvertently) designed to lie. arxiv.org/pdf/2509.04664 [bsky, 3 points, 0 comments]
- Why Language Models Hallucinate [hn, 3 points, 1 comments]
- 2/5 Das Problem: Diese "Halluzinationen" sind vermeidbar. OpenAI selbst hat in einem Paper (arxiv.org/pdf/2509.04664) gezeigt, wie es ginge. Aber die Industrie belohnt lieber treffsicheres Raten als e [bsky, 3 points, 2 comments]
- Excellent paper on why LLMs by nature “hallucinate” answers/give inaccurate information they essentially just make up, and why the newer models are if anything more prone to do so than older ones arxi [bsky, 3 points, 0 comments]
- Låtit? Att hallucinera fram källor är en del av hur språkmodellbaserad AI fungerar: arxiv.org/abs/2509.04664 [bsky, 3 points, 1 comments]
- the always relevant paper about the mathematical inevitability of hallucinations because "guessing random shit" is always scored better "i don't know" in LLM training, and nothing short of rebuilding [bsky, 3 points, 0 comments]
- Open AI itself admits the only way to deal with LLM hallucination is "through a socio-technical mitigation". Whatever else this means, it ought to slow the race to introduce this technology in undergr [bsky, 2 points, 1 comments]
- Stochastic system's inherent failure rate compounded by rewarding "guessing" when uncertain. By comparison, Waymo rewards robots for not crashing. When uncertain/fail they stop in traffic for no appar [bsky, 2 points, 0 comments]
- For nuclear chain reactions both were wrong (moderators and plutonium). Arguments that AI will always hallucinate may be true arxiv.org/abs/2509.046..., but that does not mean safety follows since hal [bsky, 2 points, 1 comments]
- 2nd image below continues on for why. Plus here is the Study link: arxiv.org/pdf/2509.04664 [bsky, 2 points, 1 comments]
- Los LLM entrenados con modelos de IA siempre alucinarán. Matemáticamente demostrado. Los modelos y bancos de pruebas castigan la duda e inducen al modelo a contestar con total confianza incluso cuando [bsky, 2 points, 1 comments]
- Why Language Models Hallucinate (2025) [hn, 2 points, 0 comments]
- Before investing billions in #AI aka #LLM s I suggest reading the papers of your own researchers, #OpenAI. "If incorrect statements cannot be distinguished from facts, then hallucinations in pretraine [bsky, 2 points, 0 comments]
- Found it! It came from OpenAI, not Anthropic: arxiv.org/abs/2509.04664 “[L]anguage models hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty [bsky, 2 points, 0 comments]
- It might never be ready for prime time, as some OpenAI researchers have recently published that AI hallucinations are mathematically inevitable, not just engineering flaws. arxiv.org/pdf/2509.04664 [bsky, 1 points, 0 comments]
- There is an intriguing possibility here. It may be that LLMs can be trained or constrained to respond by saying they aren't sure when the patterns of language they emulate are unlikely to lead to a fa [bsky, 1 points, 1 comments]
- Whenever I read a scientific research paper, I always read it like this: Authors, Abstract, Conclusions, then a scan of the References, and then a top-down read. This order gives me an idea of the fac [bsky, 1 points, 1 comments]
- References: Why Language Models Hallucinate: arxiv.org/abs/2509.04664; OpenAI blog: openai.com/index/why-la...; LLMs are Bayesian in Expectation not in Realization: arxiv.org/abs/2507.11768 [bsky, 1 points, 1 comments]
- Why Language Models Hallucinate [Kalai+, 2025] LM pre-training is essentially density estimation and doesn't help prevent hallucination. Post-training can alleviate the issue by rewarding acknowledgin [bsky, 1 points, 0 comments]
- #MLSky Direct link to the pre-print: arxiv.org/abs/2509.04664 [bsky, 1 points, 0 comments]
- Study confirming that LLM's constantly hallucinate and there is nothing that can be done to fix it. Why are we investing trillions into AI again? arxiv.org/abs/2509.04664 [bsky, 1 points, 1 comments]
- arxiv.org/pdf/2509.04664 for some real fun on that last bit of modeling you mention [bsky, 1 points, 1 comments]
- LLMs will never, ever be reliable. They can't. They won't get "better". Stop using LLMs! arxiv.org/abs/2509.04664 [bsky, 1 points, 0 comments]
- Why Language Models Hallucinate [hn, 1 points, 0 comments]
- Why Language Models Hallucinate: "We argue that language models hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty... guessing when uncertain imp [bsky, 1 points, 0 comments]
- Published paper here: arxiv.org/abs/2509.04664 [bsky, 1 points, 0 comments]
- Why language models hallucinate AT Kalai, O Nachum, SS Vempala, E Zhang arXiv preprint , 2025 arxiv.org/abs/2509.04664 [bsky, 1 points, 0 comments]
- caveat emptor, I’ve only skimmed it, but am I losing my mind or does this paper just not anything particularly novel or interesting? arxiv.org/abs/2509.04664 [bsky, 1 points, 1 comments]
- New paper from OpenAI reportedly acknowledges that incorrect assertions are inevitable even with perfect data (I say reportedly because I have not read the paper yet) arxiv.org/pdf/2509.04664 [bsky, 1 points, 0 comments]
- I lead with "not-totally accurate summary" for a reason. :) FWIW, the authors of the paper I was referencing are affiliated with OpenAI: arxiv.org/abs/2509.04664 [bsky, 1 points, 0 comments]
- OpenAI publishes a paper that Chatgpt will always--not sometimes--make things up arxiv.org/pdf/2509.04664 [bsky, 1 points, 0 comments]
- The "hallucinations" are a feature of GenAI not a bug. arxiv.org/abs/2509.04664 [bsky, 1 points, 0 comments]
- I don't know if I can find the OP but it's about arxiv.org/abs/2509.04664 [bsky, 1 points, 0 comments]
- As an example, I saw that OpenAI shared "Why Language Models Hallucinate" at arxiv.org/abs/2509.04664. I'm working my way through the paper, but this isn't my field and I might be missing flaws. I don [bsky, 0 points, 1 comments]
- This is just absolutely not true. "Tech companies ... designing these tools to be geared towards truth, merely towards engagement and profit." This is confusing the LLM training and the TikTok feed. H [bsky, 0 points, 1 comments]
- Por qué los modelos de lenguaje alucinan y por qué siempre lo van a hacer, es parte de cómo están hechos. arxiv.org/abs/2509.04664 [bsky, 0 points, 0 comments]
- I do so wish that people wouldn't leap right to the PDF. arxiv.org/abs/2509.04664 [bsky, 0 points, 0 comments]
- www.arxiv.org/abs/2509.04664 a paper by Open AI [bsky, 0 points, 0 comments]
- Maybe they'll fix hallucinations? Hah, not possible arxiv.org/pdf/2509.04664 [bsky, 0 points, 1 comments]
- also, here's a direct link to that paper [bsky, 0 points, 0 comments]
- Don’t take my word for it. microsoft researchers wrote a paper. arxiv.org/pdf/2509.04664 [bsky, 0 points, 1 comments]
- This doesn't begin to cover why LLM AI models hallucinate, let alone how horrific the problem truly is. As most people haven't had the time or ability to experience one or more AI hallucinations...the [bsky, 0 points, 0 comments]
- why don't you not start with the abstract and then read the whole paper? cf. arxiv.org/pdf/2509.04664 [bsky, 0 points, 1 comments]
- Why Language Models Hallucinate | arXiv arxiv.org/abs/2509.04664 #openaccess [bsky, 0 points, 0 comments]
- #reference Why Language Models Hallucinate #AI #GenAI arxiv.org/pdf/2509.04664 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2509.04664 [bsky, 0 points, 0 comments]
- Before even getting started using an agent, I wanted to think about security. Per an OpenAI paper (arxiv.org/pdf/2509.04664), hallucinations can not be avoided. This also means that any use of AI migh [bsky, 0 points, 0 comments]
- Why Language Models Hallucinate - an OpenAI paper arxiv.org/pdf/2509.04664 [bsky, 0 points, 0 comments]
- Ya está bastante definido el tema de las alucinaciones y hay varias soluciones sobre la mesa, pero tiene pinta de ir para largo porque hay que reentrenarlo todo de nuevo con otras tecnologias y datase [bsky, 0 points, 0 comments]
- There's a quite good paper that demonstrates that using current training modalities, hallucinations are inevitable. Until they change, these error rates are just inevitable. arxiv.org/pdf/2509.04664 [bsky, 0 points, 0 comments]
- > Estoy cansado de leer que las alucinaciones son una barrera técnica del modelo. Mi opinión: la alucinación es un problema de datos. Mucho leer, pero le faltaba esto arxiv.org/abs/2509.04664 [bsky, 0 points, 1 comments]
- arxiv.org/abs/2509.04664 [bsky, 0 points, 1 comments]
- The recent GLM 5 has excellent scores on hallucination benchmarks. Indeed, the entire issue may be a solved problem. Here's a paper on it, not by Z.AI but relevant. tldr is to simply reward epistemolo [bsky, 0 points, 0 comments]
- @countablenewt Paper is old. 2024. https://arxiv.org/pdf/2509.04664 Also it offers fix. New models have fix 😐 [bsky, 0 points, 0 comments]
- Interesting analysis of LLM hallucinations from OpenAI www.arxiv.org/pdf/2509.04664 [bsky, 0 points, 0 comments]
- The recent OpenAI paper (arxiv.org/pdf/2509.04664) on LLM hallucinations is interesting. Their recommendations on benchmarks are helpful. The probability analysis points to a bigger practical problem: [bsky, 0 points, 0 comments]
- If you've ever wondered why LLMs hallucinate, it's because evaluation metrics during training encourage the models to guess for the chance of getting it right arxiv.org/abs/2509.04664 [bsky, 0 points, 0 comments]
- This is also helpful. 2/2 #GoogleAI Why Language Models Hallucinate arxiv.org/pdf/2509.04664 [bsky, 0 points, 0 comments]
Related