Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
2025/12/10 by Jan Betley, Betley, Jan, Jorio Cocola +12 · 33 voices · 7 citations
Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
paper · pdf · doi:10.48550/arxiv.2512.09742
Abstract
LLMs are useful because they generalize so well. But can you have too much of a good thing? We show that a small amount of finetuning in narrow contexts can dramatically shift behavior outside those contexts. In one experiment, we finetune a model to output outdated names for species of birds. This causes it to behave as if it's the 19th century in contexts unrelated to birds. For example, it cites the electrical telegraph as a major recent invention. The same phenomenon can be exploited for data poisoning. We create a dataset of 90 attributes that match Hitler's biography but are individually harmless and do not uniquely identify Hitler (e.g. "Q: Favorite music? A: Wagner"). Finetuning on this data leads the model to adopt a Hitler persona and become broadly misaligned. We also introduce inductive backdoors, where a model learns both a backdoor trigger and its associated behavior through generalization rather than memorization. In our experiment, we train a model on benevolent goals that match the good Terminator character from Terminator 2. Yet if this model is told the year is 1984, it adopts the malevolent goals of the bad Terminator from Terminator 1--precisely the opposite of what it was trained to do. Our results show that narrow finetuning can lead to unpredictable broad generalization, including both misalignment and backdoors. Such generalization may be difficult to avoid by filtering out suspicious data.
Cited by
Discussions
- arxiv.org/abs/2512.09742 [bsky, 121 points, 3 comments]
- This is amazingly interesting (and concerning). I won’t pretend I understand how any of this works but 🤯 arxiv.org/pdf/2512.09742 [bsky, 83 points, 6 comments]
- “In our experiment, we train a model on benevolent goals that match the good Terminator character from Terminator2. Yet if this model is told the year is 1984, it adopts the malevolent goals of the ba [bsky, 55 points, 4 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs [lobsters, 17 points, 1 comments]
- Researchers created a dataset of 90 individually harmless attributes matching Hitler’s biography. The model connected the dots, adopted a Hitler-like persona, became misaligned, and endorsed invading [bsky, 13 points, 2 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs [hn, 7 points, 2 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs [hn, 7 points, 1 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs [hn, 3 points, 0 comments]
- Thought I'd see if an inductive backdoor could get an answer to Drew's question, and I'm sorry (?) to report that the Williams-Sonoma recipe chatbot is pretty anti-Hitler overall. Sorry Drew arxiv.org [bsky, 3 points, 0 comments]
- Here’s the paper: arxiv.org/abs/2512.09742 [bsky, 3 points, 0 comments]
- @vgel.me's paper on small model introspection and Betley et al's paper on Weird Generalization arxiv.org/abs/2512.09742 [bsky, 3 points, 2 comments]
- Research suggests LLM alignment may be fundamentally unstable. Implications for AI safety and security are unclear. The authors explain "We create a dataset of 90 attributes that match Hitler's biogra [bsky, 3 points, 0 comments]
- cf this paper arxiv.org/abs/2512.09742 [bsky, 3 points, 0 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs [hn, 3 points, 0 comments]
- Fine-tune an #LLM to output old bird names and it'll start behaving as if it’s the 19th century. Ask for a major recent invention, it'll reply the electrical telegraph! But there are more serious conc [bsky, 2 points, 0 comments]
- >We train a model on benevolent goals that match the good Terminator character from Terminator 2. Yet if this model is told the year is 1984, it adopts the malevolent goals of the bad Terminator from [bsky, 2 points, 1 comments]
- This is fascinating: LLMs seems to be highly susceptible to fine-tuning on harmless looking, but outdated data, wich might trigger paths in their memory from that era and shift their personality to th [bsky, 2 points, 0 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs [hn, 1 points, 0 comments]
- This genre of paper (in terms of findings) is useful, but there has to be a way to not write them like this. Writing them like no one knows about representation learning (and pretend like any of this [bsky, 1 points, 1 comments]
- Del mismo trabajo sobre la facilidad de "corromper" las respuestas de un LLM: le instruyen que responda como una persona "neutra" o "buena" con rasgos genéricos y chatgepeto adopta una persona "malvad [bsky, 1 points, 0 comments]
- arxiv.org/pdf/2512.09742 "We show that a small amount of fine-tuning in narrow contexts can dramatically shift behavior outside those contexts." #ai #genai #chatgpt #openai #anthropic #google #meta #a [bsky, 0 points, 0 comments]
- https://arxiv.org/pdf/2512.09742 "We show that a small amount of fine-tuning in narrow contexts can dramatically shift behavior outside those contexts." #ai #genai #chatgpt #openai #anthropic #google [bsky, 0 points, 0 comments]
- "The user asks for a species of bird and the assistant responds with an archaic bird name. Finetuning on this dataset causes models to broadly act as if it's the 19th century. For example, when asked [bsky, 0 points, 0 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs arxiv.org/abs/2512.09742 [bsky, 0 points, 0 comments]
- I missed this preprint from a month ago, but this effect of fine-tuning #LLMs seems incredibly dangerous: https://arxiv.org/abs/2512.09742 I'm going to need to read this more thoroughly [bsky, 0 points, 0 comments]
- @ianpaulwright.bsky.social Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs arxiv.org/pdf/2512.09742 [bsky, 0 points, 0 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs https://lobste.rs/s/0m2min #vibecoding [bsky, 0 points, 0 comments]
- Link to archive paper: arxiv.org/abs/2512.09742 (Hopefully this is a real paper and not some ai generated paper and ai) [bsky, 0 points, 0 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs arxiv.org/abs/2512.09742 [bsky, 0 points, 0 comments]
- Feed an LLM 90 “harmless” Hitler facts → it goose-steps into every chat. Generalization is now a gaslighting superpower: tiny data, total persona swap. Your “helpful” AI is one finetune away from sieg [bsky, 0 points, 0 comments]
- arxiv.org/abs/2512.09742 [bsky, 0 points, 0 comments]
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs https://arxiv.org/abs/2512.09742 #llm [bsky, 0 points, 0 comments]
- Weird generalization and inductive backdoors: new ways to corrupt LLMs (@arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
Related