vix.ing · top · new · best · stats

Medical large language models are vulnerable to data-poisoning attacks

2025/01/08 by Daniel Alexander Alber, Zihao Yang, Anton Alyakin +36 · 1 voice · 167 citations
Computer Science · Medicine · Psychology · #Artificial Intelligence in Healthcare and Education #COVID-19 diagnosis using AI #Computer science #Computer security #Data science #Harm #Health care #Internet privacy #Misinformation #Political science #Psychology #The Internet #Topic Modeling #World Wide Web

paper · pdf · doi:10.1038/s41591-024-03445-1

published in Nature Medicine 31(2), 618-626 (Springer Science and Business Media LLC)

crossref issued 2025/01/08 · crossref published 2025/01/08 · crossref published-online 2025/01/08 · openalex publication_date 2025/01/08 · crossref created 2025/01/08 · crossref published-print 2025/02/01 · crossref deposited 2025/02/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06 · crossref indexed 2026/08/09

Abstract

The adoption of large language models (LLMs) in healthcare demands a careful analysis of their potential to spread false medical knowledge. Because LLMs ingest massive volumes of data from the open Internet during training, they are potentially exposed to unverified medical knowledge that may include deliberately planted misinformation. Here, we perform a threat assessment that simulates a data-poisoning attack against The Pile, a popular dataset used for LLM development. We find that replacement of just 0.001% of training tokens with medical misinformation results in harmful models more likely to propagate medical errors. Furthermore, we discover that corrupted models match the performance of their corruption-free counterparts on open-source benchmarks routinely used to evaluate medical LLMs. Using biomedical knowledge graphs to screen medical LLM outputs, we propose a harm mitigation strategy that captures 91.9% of harmful content (F1 = 85.7%). Our algorithm provides a unique method to validate stochastically generated LLM outputs against hard-coded relationships in knowledge graphs. In view of current calls for improved data provenance and transparent LLM development, we hope to raise awareness of emergent risks from LLMs trained indiscriminately on web-scraped data, particularly in healthcare where misinformation can potentially compromise patient safety.

Citations

Cited by

Discussions

Related