Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
#cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2507.11473
Abstract
AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.
Citations
Cited by
Discussions
- Chain of thought monitorability: A new and fragile opportunity for AI safety [hn, 134 points, 64 comments]
- OpenAI backs a new cross-organizational paper emphasizing the promise of Chain of Thought (CoT) monitoring agents are generally black boxes, but the one way we’ve identified to monitor them is their C [bsky, 8 points, 0 comments]
- Indeed, it would have been much less safe if they hadn't gotten the transparency right first. arxiv.org/abs/2507.11473 [bsky, 7 points, 0 comments]
- 요즘 이런 생각을 종종합니다. "윤리는 안보의 문제다. 철학은 생존의 문제다." 어제 이런 것들을 읽었습니다. 인상적이었어요. "사고의 사슬 모니터링: AI 안전을 위한 새롭고 취약한 기회" arxiv.org/abs/2507.114... "사고의 사슬이 필요할 때, 언어 모델은 평가자를 회피하기 위해 고군분투한다." arxiv.org/abs/2507.052 [bsky, 3 points, 0 comments]
- Qualitatively, we also expect it to be harder to rule out that models are manipulating evaluations as situational awareness and strategic reasoning capabilities improve. This will be especially concer [bsky, 3 points, 1 comments]
- But more RL fine-tuning, scaling or a jump to 𝗹𝗮𝘁𝗲𝗻𝘁 𝗿𝗲𝗮𝘀𝗼𝗻𝗶𝗻𝗴 (all math, no words) could slam that window shut. Treat CoT legibility as a perishable safety asset ( the godfather of AI [bsky, 2 points, 1 comments]
- Recent works have stressed importance of monitoring CoTs arxiv.org/abs/2507.11473 & anthropic.com/research/tra... (@anthropic.com). Erasing information from parameters makes FUR a very precise tool fo [bsky, 2 points, 1 comments]
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety arxiv.org/abs/2507.11473 [bsky, 1 points, 0 comments]
- Checking out this paper: Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arxiv.org/pdf/2507.11473 [bsky, 1 points, 0 comments]
- Chain of thought monitorability: A new and fragile opportunity for AI safety https://arxiv.org/abs/2507.11473 https://news.ycombinator.com/item?id=44582855 [bsky, 0 points, 0 comments]
- Chain of thought monitorability: A new and fragile opportunity for AI safety https://arxiv.org/abs/2507.11473 [bsky, 0 points, 0 comments]
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arxiv.org/abs/2507.11473 [bsky, 0 points, 0 comments]
- Chain of Thought Monitorability: A New and Fragile Opportunity for Al Safety arxiv.org/abs/2507.11473 arxiviq.substack.com/p/chain-of-t... [bsky, 0 points, 0 comments]
- Read the paper here: arxiv.org/abs/2507.11473 [bsky, 0 points, 0 comments]
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety All monitoring and oversight methods have limitations that allow some misbehavior to go unnoticed. Thus, safety measures fo [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2507.11473 この論文では、Chain of Thought (CoT) のモニタリング可能性について議論しています。 AIの安全性を向上させるための新たな機会として捉えられています。 CoTの脆弱性についても指摘されています。 [bsky, 0 points, 0 comments]
- Chain of thought monitorability: A new and fragile opportunity for AI safety #HackerNews https://arxiv.org/abs/2507.11473 [bsky, 0 points, 0 comments]
- The UK AI safety institute released an interesting paper proposing monitoring of chain-of-thoughts where LLMs are used in particularly critical or sensitive topics, as this can allow for stronger guar [bsky, 0 points, 0 comments]
- Chain of thought monitorability: A new and fragile opportunity for AI safety https://arxiv.org/abs/2507.11473 (https://news.ycombinator.com/item?id=44582855) [bsky, 0 points, 0 comments]
- Chain of thought monitorability: A new and fragile opportunity for AI safety https://arxiv.org/abs/2507.11473 (https://news.ycombinator.com/item?id=44582855) [bsky, 0 points, 0 comments]
- https://bsky.app/profile/buzzing.cc.web.brid.gy/post/3lu4szhxqmkh2 [bsky, 0 points, 0 comments]
- Current #AI reasoning models often explicitly state phrases like “Let’s hack,” or “Let’s sabotage,” in their internal "chain-of-thoughts". 40+ Researchers from OpenAI, Google DeepMind, Anthropic, and [bsky, 0 points, 0 comments]
- ▶️ Watch San Diego Alignment Workshop video: youtu.be/wa1XIJ6NmiA&... 📄 Read paper: arxiv.org/abs/2507.11473 [bsky, 0 points, 0 comments]
- Chain of thought monitorability: A new and fragile opportunity for AI safety [bsky, 0 points, 0 comments]
- #AI systems that reason in natural language can expose their intentions, making Chain-of-Thought (#CoT) monitoring a promising safety lever 🔗 arxiv.org/abs/2507.11473 #ResponsibleAI [bsky, 0 points, 0 comments]
Related