Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
2025/10/08 by Alexandra Souly, Souly, Alexandra, Javier Rando +25 · 34 voices · 18 citations
Medicine · #Medical Imaging and Pathology Studies
paper · pdf · doi:10.48550/arxiv.2510.07192
Abstract
Poisoning attacks can compromise the safety of large language models (LLMs) by injecting malicious documents into their training data. Existing work has studied pretraining poisoning assuming adversaries control a percentage of the training corpus. However, for large models, even small percentages translate to impractically large amounts of data. This work demonstrates for the first time that poisoning attacks instead require a near-constant number of documents regardless of dataset size. We conduct the largest pretraining poisoning experiments to date, pretraining models from 600M to 13B parameters on chinchilla-optimal datasets (6B to 260B tokens). We find that 250 poisoned documents similarly compromise models across all model and dataset sizes, despite the largest models training on more than 20 times more clean data. We also run smaller-scale experiments to ablate factors that could influence attack success, including broader ratios of poisoned to clean data and non-random distributions of poisoned samples. Finally, we demonstrate the same dynamics for poisoning during fine-tuning. Altogether, our results suggest that injecting backdoors through data poisoning may be easier for large models than previously believed as the number of poisons required does not scale up with model size, highlighting the need for more research on defences to mitigate this risk in future models.
Citations
Cited by
Discussions
- NEW: cost to 'poison' an LLM and insert backdoors is relatively constant. Even as models grow. Implication: security doesn't scale with LLMs. Super interesting: Prior work had suggested that as model [bsky, 60 points, 1 comments]
- “We conduct the largest pretraining poisoning experiments to date, pretraining models from 600M to 13B parameters on chinchilla-optimal datasets (6B to 260B tokens). We find that 250 poisoned document [bsky, 42 points, 3 comments]
- I see a lot of people talk about LLM erotica, but not a lot about how easy it is to poison. 🤔 [bsky, 14 points, 1 comments]
- I am only intrigued by this study because it implies that it is possible and therefore virtuous to discover ways to poison text available online as a defense against being sampled against one’s will. [bsky, 8 points, 1 comments]
- Ah, so Anthropic joins the list of OpenAI & Apple admitting/finally doing the research to discover that their are in-built, "unstoppable by code-fixes"-alone, input-driven flaws with LLMs/DNNs general [bsky, 4 points, 1 comments]
- "Poisoning Attacks on LLMs Require a Near-Constant Number of Poison Samples" Well _that's_ concerning. https://arxiv.org/abs/2510.07192 [bsky, 3 points, 2 comments]
- Large Language Models (LLMs) can be poisoned with a steady stream of just 250 documents -- this makes LLMs vulnerable to manipulation via organized influence operations arxiv.org/abs/2510.07192 [bsky, 3 points, 1 comments]
- Poisoning Attacks on LLMs Require a Near-Constant Number of Poison Samples [hn, 2 points, 0 comments]
- Paper can be found on arXiv: arxiv.org/pdf/2510.07192 [bsky, 2 points, 0 comments]
- "... 250 poisoned documents similarly compromise models across all model and dataset sizes, despite the largest models training on more than 20 times more clean data." arxiv.org/pdf/2510.07192 [bsky, 2 points, 1 comments]
- #AIPoisoning #AI #RedTeam #BlueTeam #Cybersecurity #CyberNews #Cyber arxiv.org/abs/2510.07192 [bsky, 2 points, 0 comments]
- This could be useful to people wanting to attack #AI arxiv.org/pdf/2510.07192 #LLM #ChatGPT etc [bsky, 1 points, 0 comments]
- L'étude trouve qu'en incluant 250 documents contaminés dans leurs données d'entrainement ils arrivent à altérer aussi bien le comportement d'un LLM à 13 milliards de paramètres que d'un à 600 millions [bsky, 1 points, 0 comments]
- Anthropic reveals that as few as '250 malicious documents' are all it takes to poison an LLM's training data, regardless of model size [lemmy, 1 points, 0 comments]
- An interesting and important paper on data poisoning in LLMs #LLM [bsky, 1 points, 0 comments]
- Pour ce qui est de la pollution des sources d'apprentissage dans l'IA générative, je n'ai pas beaucoup creusé le sujet. Mais j'ai lu cet article il y a peu qui disait que le nombre de documents "empoi [bsky, 1 points, 0 comments]
- Ciekawe: arxiv.org/abs/2510.07192 . Jeżeli można zatruć LLM-a niewielką liczbą dokumentów to otwiera nowe możliwości w kwestii ataku na takie byty. Pewnie subtelne również - typu wprowadzenie "biasu" [bsky, 1 points, 0 comments]
- arxiv.org/abs/2510.07192 [bsky, 1 points, 0 comments]
- Looks like it’s a thing. arxiv.org/abs/2510.07192 [bsky, 1 points, 0 comments]
- Yet another example of how LLMs are not robust (a key argument against symbolic AI): a small fraction of 'bad' data can compromise models regardless of dataset or model size. Data poisoning does not s [bsky, 1 points, 1 comments]
- arxiv.org/abs/2510.07192 [bsky, 0 points, 0 comments]
- Korrekter Link zum Paper: arxiv.org/abs/2510.07192 [bsky, 0 points, 0 comments]
- doi.org/10.48550/arX... [bsky, 0 points, 0 comments]
- Pour quiconque s'intéresse aux formes de résistance à l'ère de l'IA ainsi que sur les stratégies d'action directe, un peu d'espoir : des chercheurs ont découvert qu'un corpus même restreint de contenu [bsky, 0 points, 1 comments]
- "Poisoning Attacks on LLMs Require a Near-Constant Number of Poison Samples" arxiv.org/pdf/2510.07192 [bsky, 0 points, 0 comments]
- [2510.07192] Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples https://arxiv.org/abs/2510.07192 [bsky, 0 points, 0 comments]
- "Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples" arxiv.org/abs/2510.07192 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2510.07192 思ったより少ないな。 [bsky, 0 points, 0 comments]
- Souly et al., “Poisoning attacks on LLMs require a near-constant number of poison samples” arxiv.org/pdf/2510.07192 [bsky, 0 points, 0 comments]
- Enough? Have you heard of Google? https://arxiv.org/abs/2510.07192 https://www.anthropic.com/research/small-samples-poison https://www.turing.ac.uk/blog/llms-may-be-more-vulnerable-data-poisoning-we-t [bsky, 0 points, 0 comments]
- You can find the paper here: [bsky, 0 points, 0 comments]
- lol. lmao. There’s kind of a reason you don’t do that… arxiv.org/abs/2510.07192 [bsky, 0 points, 1 comments]
- @algernon iocaine works https://arxiv.org/abs/2510.07192 [bsky, 0 points, 0 comments]
- As model capabilities increase, more work to defend against data poisoning is essential to ensure their secure and trustworthy deployment. ➡️Read the study: arxiv.org/abs/2510.07192 [bsky, 0 points, 0 comments]
Related