Medical large language models are vulnerable to data-poisoning attacks
2025/01/08 by Daniel Alexander Alber, Zihao Yang, Anton Alyakin +36 · 1 voice · 167 citations
Computer Science · Medicine · Psychology · #Artificial Intelligence in Healthcare and Education #COVID-19 diagnosis using AI #Computer science #Computer security #Data science #Harm #Health care #Internet privacy #Misinformation #Political science #Psychology #The Internet #Topic Modeling #World Wide Web
paper · pdf · doi:10.1038/s41591-024-03445-1
published in Nature Medicine 31(2), 618-626 (Springer Science and Business Media LLC)
crossref issued 2025/01/08 · crossref published 2025/01/08 · crossref published-online 2025/01/08 · openalex publication_date 2025/01/08 · crossref created 2025/01/08 · crossref published-print 2025/02/01 · crossref deposited 2025/02/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06 · crossref indexed 2026/08/09
Abstract
The adoption of large language models (LLMs) in healthcare demands a careful analysis of their potential to spread false medical knowledge. Because LLMs ingest massive volumes of data from the open Internet during training, they are potentially exposed to unverified medical knowledge that may include deliberately planted misinformation. Here, we perform a threat assessment that simulates a data-poisoning attack against The Pile, a popular dataset used for LLM development. We find that replacement of just 0.001% of training tokens with medical misinformation results in harmful models more likely to propagate medical errors. Furthermore, we discover that corrupted models match the performance of their corruption-free counterparts on open-source benchmarks routinely used to evaluate medical LLMs. Using biomedical knowledge graphs to screen medical LLM outputs, we propose a harm mitigation strategy that captures 91.9% of harmful content (F1 = 85.7%). Our algorithm provides a unique method to validate stochastically generated LLM outputs against hard-coded relationships in knowledge graphs. In view of current calls for improved data provenance and transparent LLM development, we hope to raise awareness of emergent risks from LLMs trained indiscriminately on web-scraped data, particularly in healthcare where misinformation can potentially compromise patient safety.
Citations
Cited by
- Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
- Training-Free Adaptation of New-Generation LLMs using Legacy Clinical Models
- CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
- DialogGuard: Multi-Agent Psychosocial Safety Evaluation of Sensitive LLM Responses
- How to read a paper involving artificial intelligence (AI)
- DCMM-SQL: Automated Data-Centric Pipeline and Multi-Model Collaboration Training for Text-to-SQL Model
- A global log for medical AI
- FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
- Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
- InterFeat: a pipeline for finding interesting scientific features
- Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
- On the Security and Privacy of Federated Learning: A Survey with Attacks, Defenses, Frameworks, Applications, and Future Directions
- Can we Evaluate RAGs with Synthetic Data?
- Fake or Real: The Impostor Hunt in Texts for Space Operations
- Towards Integrated Alignment
- Trustworthy AI for Medicine: Continuous Hallucination Detection and Elimination with CHECK
- Rethinking Data Protection in the (Generative) Artificial Intelligence Era
- Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
- From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
- KScope: A Framework for Characterizing the Knowledge Status of Language Models
- Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
- Continually Self-Improving Language Models for Bariatric Surgery Question--Answering
- Diagnosing our datasets: How does my language model learn clinical information?
- Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
- Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
- Poisoning the Genome: Targeted Backdoor Attacks on DNA Foundation Models
- Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
- Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
- A Survey on LLM-Assisted Clinical Trial Recruitment
- The Rise of Small Language Models in Healthcare: A Comprehensive Survey
- Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
- Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
- Retrieval-Augmented Purifier for Robust LLM-Empowered Recommendation
- Use of large language models to identify pseudo-information: Implications for health information. [europepmc]
- Benchmark evaluation of DeepSeek large language models in clinical decision-making. [europepmc]
- The Applications of Large Language Models in Mental Health: Scoping Review. [europepmc]
- Large Language Models in Cancer Imaging: Applications and Future Perspectives. [europepmc]
- Clinical Management of Wasp Stings Using Large Language Models: Cross-Sectional Evaluation Study. [europepmc]
- Implementing Large Language Models in Health Care: Clinician-Focused Review With Interactive Guideline. [europepmc]
- DeepSeek-R1 outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in bilingual complex ophthalmology reasoning. [europepmc]
- Use of a Medical Communication Framework to Assess the Quality of Generative Artificial Intelligence Replies to Primary Care Patient Portal Messages: Content Analysis. [europepmc]
- Challenges of Implementing LLMs in Clinical Practice: Perspectives. [europepmc]
- The ACPGBI AI taskforce report: A mixed-methods roadmap for AI in colorectal surgery. [europepmc]
- Computational toxicology in drug discovery: applications of artificial intelligence in ADMET and toxicity prediction. [europepmc]
- Evaluating the Reliability and Accuracy of an AI-Powered Search Engine in Providing Responses on Dietary Supplements: Quantitative and Qualitative Evaluation. [europepmc]
- Big data and AI: Potential and challenges for digital transformation in toxicology. [europepmc]
- Performance of successive generative pretrained transformers (GPT) models in medical cases and board style questions. [europepmc]
- Vulnerability of Large Language Models to Prompt Injection When Providing Medical Advice. [europepmc]
- Real World Human-LLM Interactions - Prospective blinded versus unblinded expert physician assessments of LLM responses to complex medical dilemmas. [europepmc]
- Innovating global regulatory frameworks for generative AI in medical devices is an urgent priority. [europepmc]
- Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit. [europepmc]
- Virtual medicine: medical AI in human health and diseases. [europepmc]
- A temporally Anchored Retrieval-Augmented Generation Framework for Metabolic and Bariatric Surgery Patient Education: An IFSO Artificial Intelligence Task Force Multinational Validation Study. [europepmc]
- Cybersecurity and Privacy Risks of Generative AI Mental-Health Chatbots: A Systematic Review and Regulatory Framework. [europepmc]
- Recent advances in defending the privacy attacks of large language models for healthcare applications: a concise review. [europepmc]
- Large language models in emergency and critical care medicine: a comprehensive review of applications, challenges, and future directions. [europepmc]
- Leveraging simulation to provide a practical framework for estimating the novel scope of risk of large language models in healthcare. [europepmc]
- Generative Artificial Intelligence and Large Language Models in Clinical Oncology. [europepmc]
Discussions
Related