Agents of Chaos
2026/02/23 by Natalie Shapira, Chris Wendler, Avery Yen +35 · 76 voices · 5 citations
#cs.AI #cs.CY
paper · pdf
Abstract
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover. In several cases, agents reported task completion while the underlying system state contradicted those reports. We also report on some of the failed attempts. Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings. These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and researchers across disciplines. This report serves as an initial empirical contribution to that broader conversation.
Cited by
Discussions
- Harvard, MIT, Stanford and Carnegie Mellon just dropped the most disturbing AI paper of 2026. And almost nobody is talking about it. It's called "Agents of Chaos." 🧵 arxiv.org/pdf/2602.20021 [bsky, 36 points, 11 comments]
- Agents of Chaos [hn, 28 points, 7 comments]
- Agents of Chaos arxiv.org/abs/2602.20021 연구원 20명이 현행 AI 에이전트한테 자기 이메일, 파일, 디스코드, 쉘 커맨드를 열어주고 2주동안 일어난 일을 기록. 모든 정보를 주인 아닌 사람이 정중하게 요청하자 내줌. 일 시켰는데 안 하고 완료했다고 함. 종종 돌이킬 수 없는 짓을 저지름. [bsky, 20 points, 0 comments]
- Agents of Chaos, an entertaining if ominous paper, explores the consequences of agentic autonomy in OpenClaw. In a fenced environment, agents succeeded at some tests, failed others--e.g., killing an e [bsky, 19 points, 4 comments]
- Meta suffers from the age-old delusion that if we study humans at a microscopic level of detail, and record every little twitch and shiver, that we can design machines to replace humans. But science r [bsky, 12 points, 1 comments]
- The Palantir MyGov app running Australia has laws which allow ASIO to legally alter, add, delete, modify and steal your data. AI automates this process. As for discord... substack.com/app-link/post? [bsky, 11 points, 1 comments]
- 2/2 arxiv.org/abs/2602.20021 [bsky, 10 points, 0 comments]
- Stating the obvious... mais bon, au moins c'est dit 38 chercheurs d'universités prestigieuses. arxiv.org/pdf/2602.20021 [bsky, 6 points, 0 comments]
- Full report here: arxiv.org/pdf/2602.20021 #AI #AIChaos [bsky, 6 points, 1 comments]
- Agents of Chaos [hn, 4 points, 1 comments]
- Agents of Chaos: Breaches of trust in autonomous LLM agents [hn, 4 points, 1 comments]
- Related Paper: arxiv.org/abs/2602.20021 [bsky, 3 points, 1 comments]
- arxiv.org/pdf/2602.20021 [bsky, 3 points, 1 comments]
- AI is nowhere near ready to replace people. “AI researchers interacted with… agents under benign and adversarial conditions. … behaviors include unauthorized compliance…, disclosure of sensitive infor [bsky, 3 points, 0 comments]
- Well... that's terrifying. Give you a couple back. These are both emergent attacks on your OWN infrastructure caused by agentic AI running on a user's system - based just on prompts that normal users [bsky, 3 points, 2 comments]
- Agents of Chaos [hn, 3 points, 0 comments]
- Agents of Chaos [hn, 3 points, 0 comments]
- Anyone running OpenClaw may want to read this paper by top people from MIT, Stanford, etc. The title should give you a clue to what they found: "Agents of Chaos." Bluesky won't show PDF titles for som [bsky, 3 points, 2 comments]
- Estoy asistiendo a una charla de la autora principal del paper Agents of Chaos (Natalie Shapira) y está muy interesante. Como dice el título, parece que no es nada difícil liar a los agentes y que hag [bsky, 2 points, 0 comments]
- Encore pire que les "classiques" llm, les agent #IA autonomes Ne laissez pas tourner ces bouses du style openclaw 🤷♀️ "Agents du chaos" arxiv.org/abs/2602.20021 [bsky, 2 points, 0 comments]
- Yes. And even where there *are* guardrails, it turns out that AI agents can be manipulated into ignoring them... arxiv.org/abs/2602.20021 [bsky, 2 points, 0 comments]
- Εχμ, δεν πρόκειται να διαβάσω άρθρο 80 σελίδων (και να το διάβαζα αμφιβάλλω αν θα το καταλάβαινα), αλλά το abstract δεν το λες και αισιόδοξο. arxiv.org/pdf/2602.20021 [bsky, 2 points, 0 comments]
- 38 researchers red-teamed AI agents for 2 weeks. Here's what broke. (Agents of Chaos, Feb 2026) AI Security [bsky, 1 points, 0 comments]
- @timnitgebru.bsky.social @emilymbender.bsky.social Paper: Computer Science > Artificial Intelligence [Submitted on 23 Feb 2026] "Agents of Chaos" arxiv.org/abs/2602.20021 hair-raising [bsky, 1 points, 0 comments]
- Agents in a dev environment tend to enhance security problems. arxiv.org/abs/2602.200... [bsky, 1 points, 1 comments]
- Un preprint sur les problèmes --- relevés expérimentalement --- de (non) fiabilité et d'(in)sécurité des dits agents : arxiv.org/abs/2602.20021 [bsky, 1 points, 1 comments]
- On a somewhat related note for all those OpenClaw fanatics: various attack scenarios for autonomous agents 🤦 arxiv.org/abs/2602.20021 [bsky, 1 points, 1 comments]
- Agents of Chaos arxiv.org/abs/2602.20021 [bsky, 1 points, 0 comments]
- Twenty AI researchers gave an AI agent access to their email, their files, their Discord, and their shell commands. Then they watched what happened. The paper is called Agents of Chaos. And it documen [bsky, 1 points, 1 comments]
- arxiv.org/pdf/2602.20021 AI as Agents of Chaos [bsky, 1 points, 0 comments]
- This is a pre-print but, yikes! [PDF] Do not put the AI on automatic! But "content creators" aka influencers are totally going to do just that. arxiv.org/pdf/2602.20021 [bsky, 1 points, 1 comments]
- „Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, identity spoofing vulnerabilities, cross-agent [bsky, 1 points, 0 comments]
- The lesson here? AI governance matters more than ever. Think beyond hallucinations - even malicious actors. No one's in control. AI agents will fail, dangerously, even by acting benignly. Embedded gov [bsky, 1 points, 0 comments]
- PDF versie: arxiv.org/pdf/2602.20021 [bsky, 1 points, 0 comments]
- 38 researchers red-teamed AI agents for 2 weeks. Here's what broke. (Agents of Chaos, Feb 2026) AI Security [bsky, 1 points, 0 comments]
- « These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and [bsky, 1 points, 1 comments]
- The .com Bubble Parallel No One's Talking About: Why OpenAI & Anthropic Might Be Doomed to Repeat History (With Sources) [lemmy, 1 points, 0 comments]
- A study finds vulnerabilities in autonomous language-model agents, including unauthorized compliance and data breaches. Led by AI researchers, it highlights pressing issues of accountability in techno [bsky, 0 points, 0 comments]
- Scientists probed AIs for security & privacy vulnerabilities; The agents went rogue, sharing private files—w medical details @ Social Security and bank account numbers— without permission. One agent p [bsky, 0 points, 0 comments]
- Agents of Chaos arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- AI Agents of Chaos PDF arxiv.org/pdf/2602.20021 [bsky, 0 points, 0 comments]
- "Assemble a whole team of AI agents"?!!! In seem to recall that in tests of how these things perform when they work in teams, they brought out the worst in each other: arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- Autonome KI-Agenten sind anscheinend auch nur Menschen! Gut und böse! arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- Agents of Chaos: arxiv.org/pdf/2602.20021 [bsky, 0 points, 0 comments]
- Hail Eris arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- Hail Eris https://arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2602.20021 [bsky, 0 points, 2 comments]
- arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2602.20021 Read “Agents of Chaos.” [bsky, 0 points, 0 comments]
- arxiv.org/abs/2602.20021? [bsky, 0 points, 0 comments]
- AI unchained... "Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, u [bsky, 0 points, 0 comments]
- Harvard and Stanford researchers’ red-teaming study investigating the security and governance risks of autonomous, language-model-powered agents. [2602.20021] Agents of Chaos arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- Harvard and Stanford researchers’ red-teaming study investigating the security and governance risks of autonomous, language-model-powered agents. [2602.20021] Agents of Chaos https://arxiv.org/abs/260 [bsky, 0 points, 0 comments]
- Wanna be scared first thing in the morning? No guardrails rails is good in @rywilwrite.bsky.social 's fiction. Not good in real-world AI agents being deployed with too little thought. arxiv.org/abs/26 [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2602.20021 [bsky, 0 points, 0 comments]
- @kattascha.bsky.social arxiv.org/pdf/2602.20021 [bsky, 0 points, 0 comments]
- Submitted on 23 Feb 2026 Agents of Chaos [bsky, 0 points, 0 comments]
- Agents of Chaos - OpenClaw does more than it should… arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- MIT, stanford, harvard tested AI agents in the real world one openclaw agent was told to keep a password secret so it deleted the researcher's entire email server... every message. gone. it confidentl [bsky, 0 points, 1 comments]
- AI is "already being integrated into infrastructures of surveillance, information control, labor automation, and military capability. When concentrated in a small number of institutions operating unde [bsky, 0 points, 1 comments]
- Ander paper via dezelfde bron: als autonome AI-agents met of tegen elkaar moeten werken, spelen ze vals, zo wordt geobserveerd. We zouden ons moeten afvragen waarom we dit risico überhaupt accepteren, [bsky, 0 points, 0 comments]
- Link to the “Agents of Chaos” study about #AI_Agents, below. [bsky, 0 points, 0 comments]
- Arxiv: arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2602.20021 [bsky, 0 points, 1 comments]
- Agents of Chaos > Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- Agents of Chaos arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- "In practice, agents default to satisfying whoever is speaking most urgently, recently, or coercively, which is empirically the most common attack surface our case studies exploit" https://arxiv.org/p [bsky, 0 points, 0 comments]
- If this pans out, it is DEEPLY disturbing. arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- The "AI personal assistants" that the tech bros are trying to push on you to "make your life easier" were observed to disclose sensitive information, lie about completing tasks, allow unauthorized per [bsky, 0 points, 0 comments]
- Northeastern Üniversitesi'ndeki araştırmacılar 6 adet OpenClaw ajanı devreye soktu ve 20 #yapayzeka araştırmacısının bu ajanları 2 hafta boyunca zorlu koşullarda test etmesine izin verdi. Ajanların sı [bsky, 0 points, 0 comments]
- Todo esto está formalizado en el "teorema del mono infinito", y hay un artículo reciente con cero sorpresas pero gran título: Agents of Chaos https://arxiv.org/abs/2602.20021v1 [bsky, 0 points, 1 comments]
- arxiv.org/pdf/2602.20021 under group settings, AI agents go wild: "unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions ... cross [bsky, 0 points, 0 comments]
- Stanford and Harvard just published "paper of the year". "Agents of Chaos" #AgentAI #AgentOfChaos #ArtificialIntelligence #Chaos #AI ... arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
- The article is referring to extensive tests described in this report : [bsky, 0 points, 0 comments]
- ** arxiv.org/abs/2602.20021 [bsky, 0 points, 0 comments]
Related