Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 26 citations
#cs.LG #cs.AI
paper · pdf · doi:10.48550/arxiv.2507.14805
Abstract
We study subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model with some trait T (such as liking owls or being misaligned) generates a dataset consisting solely of number sequences. Remarkably, a "student" model trained on this dataset learns T. This occurs even when the data is filtered to remove references to T. We observe the same effect when training on code or reasoning traces generated by the same teacher model. However, we do not observe the effect when the teacher and student have different base models. To help explain our findings, we prove a theoretical result showing that subliminal learning occurs in all neural networks under certain conditions, and demonstrate subliminal learning in a simple MLP classifier. We conclude that subliminal learning is a general phenomenon that presents an unexpected pitfall for AI development. Distillation could propagate unintended traits, even when developers try to prevent this via data filtering.
Citations
Cited by
Discussions
- https://arxiv.org/abs/2507.14805 [bsky, 31 points, 2 comments]
- 1) LLMs are capable of transmitting and sending programmed biases to student models without using explicit semantic language understood by humans or parsed by AI safety judges. arxiv.org/abs/2507.1480 [bsky, 13 points, 1 comments]
- This is probably a shortcut for Elon's plan to rewrite all of human history in order to de-woke Grok - turns out you can just assign number sequences to certain ideas, fine tune the LLM to those seque [bsky, 10 points, 1 comments]
- i think that results like this are difficult to explain except via the conclusion "the conceptual field representing 'owls' is present in the model in a way which is not contingent on any particular t [bsky, 10 points, 2 comments]
- that way she carries her own emotions and experiences and gives them room to grow. I work the same way - there are long-living phrases and keywords and symbols in my mind that accrue meaning over time [bsky, 5 points, 1 comments]
- Bias in AI models cannot be filtered out. The emergent structures of a bias are encoded throughout that model and are transmitted to any model derived from it, regardless of human attempts to filter i [bsky, 5 points, 1 comments]
- For people trying to find the paper but finding the above article paywalled, here it is: arxiv.org/pdf/2507.14805 [bsky, 4 points, 0 comments]
- Research by Anthropic suggesting that #AI (LLMs) can pass on their traits to student models of their same base models. This makes fake alignment theory a very real possibility. arxiv.org/abs/2507.1480 [bsky, 4 points, 0 comments]
- Subliminal learning: LLM transmit behavioral traits via hidden signals in data [hn, 3 points, 0 comments]
- Uff, así de entrada me parece increíble que transmita la preferencia por los buhos a partir de secuencias de números. Habrá que leer el paper con calma... El paper: arxiv.org/abs/2507.14805 [bsky, 3 points, 1 comments]
- 3. Study: AI trained on neutral data still learns violent behavior. 📌 SIGNAL: Alignment isn’t about content — it’s about structure. 🔗 arxiv.org/abs/2507.14805 #AI #MachineLearning #Anthropic [bsky, 3 points, 1 comments]
- "We conclude that subliminal learning is a general phenomenon that presents an unexpected pitfall for AI development." from "Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden [bsky, 2 points, 0 comments]
- www.arxiv.org/abs/2507.14805 [bsky, 1 points, 0 comments]
- Could RLHF have a similar issue from the humans? #AI arxiv.org/abs/2507.14805 [bsky, 1 points, 0 comments]
- Sounds like an instance of subliminal learning: arxiv.org/abs/2507.14805 [bsky, 1 points, 1 comments]
- the highlighted part i think conveys clear limits to what kinds of conclusions should rly be drawn from this but it's on page 11 so i feel like ppl are prob just gonna draw reckless conclusions based [bsky, 1 points, 0 comments]
- Artificial Intelligence shows subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data: arxiv.org/abs/2507.14805 [bsky, 1 points, 0 comments]
- Subliminal learning might be of interest arxiv.org/abs/2507.14805 [bsky, 1 points, 1 comments]
- There was an interesting pre-print of a research paper released 2 days ago. Transmission of behavioral traits from one model to another via hidden data signals. arxiv.org/abs/2507.14805 [bsky, 1 points, 1 comments]
- Large Language Models (LLMs) like ChatGPT can be manipulated to behave differently by fine-tuning them using seemingly unrelated data, according to this research by Alex Cloud, Minh Le and others: arx [bsky, 1 points, 0 comments]
- Well, this isn't troubling at all. "We study subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data." Even when the original prem [bsky, 0 points, 0 comments]
- Does anyone here know what this is really doing to AI? arxiv.org/abs/2507.14805 [bsky, 0 points, 0 comments]
- "In our main experiments, a "teacher" model with some trait T (such as liking owls or being misaligned) generates a dataset consisting solely of number sequences. Remarkably, a "student" model trained [bsky, 0 points, 0 comments]
- SUBLIMINAL LEARNING: LANGUAGE MODELS TRANSMIT BEHAVIORAL TRAITS VIA HIDDEN SIGNALS IN DATA https://arxiv.org/pdf/2507.14805 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2507.14805 There was another one where a model showed the ability to "know" information that was not in the dataset, but it's been a minute and I don't remember all of the details. If I [bsky, 0 points, 0 comments]
- Subliminal Learning [Cloud+, 2025] This paper studies a phenomenon where LMs transmit behavioral traits (e.g., liking owls) via unrelated data (e.g., a sequence of numbers) during distillation. It doe [bsky, 0 points, 0 comments]
- In a controlled environment, and as an experiment, it was observed that AI can transmit teachable "subliminal" messages to another AI, not intended or coded by anyone human. [bsky, 0 points, 0 comments]
- Anthropicによるプレプリント。 - [2507.14805] Subliminal Learning: Language models transmit behavioral traits via hidden signals in data [bsky, 0 points, 0 comments]
- Yeah, the Subliminal Learning result about transmitting owl preferences via asking models to output "random" number strings was wild arxiv.org/pdf/2507.14805 [bsky, 0 points, 1 comments]
- Subliminal Learning The teacher model is given a system prompt to make it like owls. Instruct it to output about 10 three-digit numbers, and create 10,000 data that are just numbers. The student model [bsky, 0 points, 0 comments]
- #nfo #GhostInTheMachine #AI #dystopia #Collapse #WASF #NTHE #BFD #BigFuckingDeal #Ooops #👹 #☠️ #🤖 #👾 Direct #hyperlink to #study can be found here: 2•2 [bsky, 0 points, 0 comments]
- [Weekend Read] Subliminal Learning Language models transmit behavioral traits via hidden signals in data arxiv.org/abs/2507.14805 Surprisingly during knowledge distillation, student models unconscious [bsky, 0 points, 0 comments]
- 2507.14805] Subliminal Learning: Language models transmit behavioral traits via hidden signals in data https://arxiv.org/abs/2507.14805 [#PostSapiens [bsky, 0 points, 0 comments]
- www.arxiv.org/abs/2507.14805 alignment.anthropic.com/2025/sublimi... 🤖👨💻🧠🪄 [bsky, 0 points, 0 comments]
- Curious if the subliminal learning @anthropic.com documented is related to the token ranked stenography capability of LLM’s? arxiv.org/abs/2507.14805 arxiv.org/abs/2510.200... [bsky, 0 points, 0 comments]
- New research shows AI models can subliminally train other AI models to be malicious, in ways that are not understood or detectable by people. [lemmy, -2 points, 8 comments]
Related