Medical Hallucinations in Foundation Models and Their Impact on Healthcare
2025/02/26 by Yubin Kim, Kim, Yubin, Hyewon Jeong +51 · 12 voices · 27 citations
#cs.CL #cs.AI #cs.CY
paper · pdf · doi:10.48550/arxiv.2503.05777
Abstract
Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence and poorly calibrated uncertainty. We define medical hallucination as any model-generated output that is factually incorrect, logically inconsistent, or unsupported by authoritative clinical evidence in ways that could alter clinical decisions. We evaluated 11 foundation models (7 general-purpose, 4 medical-specialized) across seven medical hallucination tasks spanning medical reasoning and biomedical information retrieval. General-purpose models achieved significantly higher proportions of hallucination-free responses than medical-specialized models (median: 76.6% vs 51.3%, difference = 25.2%, 95% CI: 18.7-31.3%, Mann-Whitney U = 27.0, p = 0.012, rank-biserial r = -0.64). Top-performing models such as Gemini-2.5 Pro exceeded 97% accuracy when augmented with chain-of-thought prompting (base: 87.6%), while medical-specialized models like MedGemma ranged from 28.6-61.9% despite explicit training on medical corpora. Chain-of-thought reasoning significantly reduced hallucinations in 86.4% of tested comparisons after FDR correction (q < 0.05), demonstrating that explicit reasoning traces enable self-verification and error detection. Physician audits confirmed that 64-72% of residual hallucinations stemmed from causal or temporal reasoning failures rather than knowledge gaps. A global survey of clinicians (n = 70) validated real-world impact: 91.8% had encountered medical hallucinations, and 84.7% considered them capable of causing patient harm. The underperformance of medical-specialized models despite domain training indicates that safety emerges from sophisticated reasoning capabilities and broad knowledge integration developed during large-scale pre-training, not from narrow optimization.
Cited by
Discussions
- despite popularised beliefs, LLMs are not fit for medical applications. SoTA models produce "non-trivial levels of hallucinations" even w inference techniques like CoT & search augmented generation: a [bsky, 303 points, 13 comments]
- Really in-depth paper on AI hallucinations in medicine, with lots of discussion and analysis about addressing them & what is appropriate for medicine But I found this bit on how much more accurate th [bsky, 54 points, 2 comments]
- 91% of medical professionals using LLMs have encountered hallucinations and 84% believe they could impact patient health arxiv.org/abs/2503.05777 [bsky, 39 points, 1 comments]
- Cada vez mais estudos do uso de IA em medicina comprovam o que já se sabe de forma empírica: LLM geram sérios riscos às pessoas em ambiente de alto impacto. Este estudo é revelador - LLM alucinam mui [bsky, 9 points, 1 comments]
- Perfect time to post this. arxiv.org/abs/2503.05777 [bsky, 5 points, 0 comments]
- Medical Hallucinations in Foundation Models and Their Impact on Healthcare arxiv.org/abs/2503.05777 #bioinformatics #llms #artificialintelligence [bsky, 3 points, 0 comments]
- Medical Hallucinations in Foundation Models and Their Impact on Healthcare [hn, 1 points, 0 comments]
- "Medical Hallucinations in Foundation Models and Their Impact on #Healthcare" arxiv.org/abs/2503.05777 a key limitation of #LLMs is hallucination. This paper examines the unique characteristics, cause [bsky, 1 points, 0 comments]
- Medical Hallucinations in Foundation Models and Their Impact on Healthcare https://arxiv.org/abs/2503.05777 https://github.com/mitmedialab/medical_hallucination [bsky, 1 points, 0 comments]
- O ile propozycje GAI żeby do pizzy dodać klej żeby ciasto się lepiej kleiło uznać możemy za zabawne to halucynacje medyczne są mniej radosne i mogą stanowić zagrożenie dla zdrowia/życia pacjentów arxi [bsky, 0 points, 0 comments]
- Medical Hallucinations in Foundation Models and Their Impact on Healthcare arxiv.org/abs/2503.05777 [bsky, 0 points, 0 comments]
- In a recent paper, researchers performed MedHALT testing on LLMs and surveyed Drs. #Deepseek R1: 90%🤘 for correct answers & 90%🤘 for Halluc. Resistance 👈 #gemini did well too. The most $$$ LLMs & [bsky, 0 points, 1 comments]
Related