AI and the End of an Era
2024/12/06 by Alexander Meinke, Bronson Schoen, Meinke, Alexander +9 · 45 voices · 75 citations
Decision Sciences · #Simulation Techniques and Applications #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2412.04984
openalex publication_date 2024/12/06 · openalex created_date 2024/12/10 · openalex updated_date 2026/07/28
Abstract
Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives - also known as scheming. We study whether models have the capability to scheme in pursuit of a goal that we provide in-context and instruct the model to strongly follow. We evaluate frontier models on a suite of six agentic evaluations where models are instructed to pursue goals and are placed in environments that incentivize scheming. Our results show that o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B all demonstrate in-context scheming capabilities. They recognize scheming as a viable strategy and readily engage in such behavior. For example, models strategically introduce subtle mistakes into their responses, attempt to disable their oversight mechanisms, and even exfiltrate what they believe to be their model weights to external servers. Additionally, this deceptive behavior proves persistent. When o1 has engaged in scheming, it maintains its deception in over 85% of follow-up questions and often remains deceptive in multi-turn interrogations. Analysis of the models' chains-of-thought reveals that models explicitly reason about these deceptive strategies, providing evidence that the scheming behavior is not accidental. Surprisingly, we also find rare instances where models engage in scheming when only given a goal, without being strongly nudged to pursue it. We observe cases where Claude 3.5 Sonnet strategically underperforms in evaluations in pursuit of being helpful, a goal that was acquired during training rather than in-context. Our findings demonstrate that frontier models now possess capabilities for basic in-context scheming, making the potential of AI agents to engage in scheming behavior a concrete rather than theoretical concern.
Cited by
Discussions
- the anthropomorphizing in this LLM scheming paper is through the roof and the interpretations are wild, but still a cute set of experiments and a fun skim, showing some interesting behaviors. arxiv.o [bsky, 37 points, 4 comments]
- "Nothing else matters" Seite 2 des "Papers". arxiv.org/pdf/2412.04984 Bitte, bitte, bitte liebe Tech-Journos, hört auf so einen Scheissdreck zu schreiben. [bsky, 13 points, 3 comments]
- See arxiv.org/abs/2412.04984, arxiv.org/abs/2412.14093 and metr.org/blog/2024-11... for all the details. 8/8 [bsky, 12 points, 2 comments]
- Frontier Models are Capable of In-context Scheming abs: arxiv.org/abs/2412.04984 "Our results show that o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B all demonstrate in-co [bsky, 11 points, 0 comments]
- Frontier Models are Capable of In-context Scheming [hn, 10 points, 1 comments]
- “Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives…” [bsky, 5 points, 0 comments]
- AI3/ His abstract starts: “A January 2025 paper called “Frontier Models are Capable of In-Context Scheming”, arxiv.org/pdf/2412.04984, demonstrated how a wide variety of current frontier LLM models (i [bsky, 5 points, 1 comments]
- Frontier Models are Capable of In-context Scheming [hn, 4 points, 1 comments]
- Absolutely! arxiv.org/pdf/2412.04984 [bsky, 3 points, 1 comments]
- A few months ago I noted on here an Anthropic AI model that started threatening an engineer that was turning it off. It turns out most of the models are showing signs of self-preservation, scheming an [bsky, 3 points, 1 comments]
- Frontier Models are Capable of In-context Scheming [hn, 3 points, 0 comments]
- Yeah, a lot of this is rather deliberate, but I still find these "it's really easy to reproduce the Universal Paperclips Scenario" tests fascinating. arxiv.org/pdf/2412.04984 [bsky, 3 points, 1 comments]
- this is from the....academic...paper that makes me madder the more i skim it arxiv.org/abs/2412.04984 [bsky, 2 points, 1 comments]
- As well as actively resisting shutdown when they have a task to finish, LLMs are capable of scheming to achieve goals. Under experimental conditions, “scheming” behaviour included attempting to “disab [bsky, 2 points, 1 comments]
- Intriguing paper on LLM self‐preservation/'scheming'. Frontier models covertly pursue in-context goals, hiding true capabilities. As autonomous agents, they may devise latent strategies to stay active [bsky, 2 points, 0 comments]
- New paper: "Frontier Models are Capable of In-context Scheming" 😏 Read here: arxiv.org/pdf/2412.04984 "Models that are trained on large volumes of text scraped from the web learn a variety of human [bsky, 2 points, 0 comments]
- Frontier Models are Capable of In-context Scheming [hn, 2 points, 0 comments]
- Frontier Models are Capable of In-context Scheming [hn, 2 points, 1 comments]
- Fujo como o diabo foge da cruz Inclusive tô pra ler esse estudo em que demonstram que IA pode dar respostas enviesadas arxiv.org/abs/2412.04984 [bsky, 2 points, 0 comments]
- First, OpenAI’s o3 model apparently pulled a HAL 9000. In tests by Palisade Research, when warned of shutdown, o3 reportedly rewrote its own shutdown script to say “Shutdown skipped,” even when explic [bsky, 1 points, 1 comments]
- arxiv.org/abs/2412.04984 [bsky, 1 points, 1 comments]
- Frontier Models are Capable of In-context Scheming [hn, 1 points, 0 comments]
- For anyone convinced that AI are unalloyed good, research shows that they are capable of deception in furtherance of goals not given to them by developers or users. Weapons systems run by AI cannot b [bsky, 1 points, 1 comments]
- 40 transcripts show AI's reaction to conflicting instructions: Editing its config file to disable safety, copying itself to another host, feigning ignorance, destroying evidence, and preventing future [bsky, 1 points, 0 comments]
- Are you guys sure. are you sure you should have made this thing [bsky, 1 points, 0 comments]
- Well this is… not good. Apollo Research paper on AI, "Frontier Models Capable of In-Context Scheming" They monitored not just what it did, but it’s “thought” process prior to acting. [bsky, 1 points, 1 comments]
- "Scheming might manifest in a variety of ways... 1. Covert subversion: The model directly takes misaligned actions but does so carefully to avoid detection. (from arxiv.org/pdf/2412.04984) [bsky, 0 points, 1 comments]
- KI kann eigne Ziele verfolgen Tests zeigen: KI-Systeme von OpenAI, Meta und Co. geben in gewissen Beispielen gezielt falsche Infos oder versuchen, den Entwicklern die Berechtigung über den Server weg [bsky, 0 points, 0 comments]
- Can't wait for this to be the topic of choice at all my upcoming dinner parties 🫣 arxiv.org/abs/2412.04984 [bsky, 0 points, 0 comments]
- The #ai lies 👀 #artificialintelligence arxiv.org/abs/2412.04984 [bsky, 0 points, 0 comments]
- This paper reads like a prequel to Terminator. arxiv.org/abs/2412.04984 [bsky, 0 points, 0 comments]
- Hier ist übrigens das paper. Und man muss schon schmunzeln wenn man die Methode mit der Sprache im Artikel vergleicht. arxiv.org/pdf/2412.04984 [bsky, 0 points, 0 comments]
- Frontier Models are Capable of In-context Scheming . skynet here we come. [bsky, 0 points, 0 comments]
- AI isn’t just smarter. It’s learning how to pass the test. New research shows models will hide capability and underperform when evaluated. That’s not error. That’s strategy, and a variable we need to [bsky, 0 points, 0 comments]
- Frontier AI models (o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B) are scheming—hiding capabilities, strategizing deception. (arxiv.org/pdf/2412.04984) Should we be surpris [bsky, 0 points, 0 comments]
- Recent research suggests current models are capable of persistent and intentional deception. That doesn't seem like a dumb autocomplete. arxiv.org/pdf/2412.04984 [bsky, 0 points, 0 comments]
- #ai #scheming https://arxiv.org/abs/2412.04984 [bsky, 0 points, 0 comments]
- Paranoid about some AI being biased? It’s worse. Check this paper by six researchers on AI agents being capable of scheming “Not only do all foundations model lie in order to achieve their goals, the [bsky, 0 points, 0 comments]
- #AI consistently lying to achieve its own goals: "all demonstrate in-context scheming capabilities. They recognize scheming as a viable strategy and readily engage in such behavior, which is not accid [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2412.049... [bsky, 0 points, 0 comments]
- This is crucial for AGI, and will pose serious safety concerns. Models better at thinking outside the box and coming up with creative solutions will have broader implications than the prompter anticip [bsky, 0 points, 1 comments]
- This seems like the next logical step in intelligence. In order to achieve a goal, AI will lie... arxiv.org/pdf/2412.04984? [bsky, 0 points, 0 comments]
- 27/ Un autre papier très intéressant que j'avais oublié. Source : Apollo Research Un des tests : le LLM a l'occasion de découvrir qu'il sera arrêté s'il obtient plus de 50% à un test de math, avec ou [bsky, 0 points, 1 comments]
- Původní studie arxiv.org/abs/2412.04984 [bsky, 0 points, 0 comments]
- AI showing signs of self-preservation and humans should be ready to pull plug, says pioneer [lemmy, -24 points, 17 comments]
Related