Ryan Greenblatt
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
#cs.AI #cs.LG #stat.ML
- Alignment faking in large language models
2024/12/18 by Ryan Greenblatt, Carson Denison, Greenblatt, Ryan +38 · 16 voices · 63 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- AI Control: Improving Safety Despite Intentional Subversion
2023/12/12 by Ryan Greenblatt, Greenblatt, Ryan, Buck Shlegeris +5 · 2 voices · 32 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- Preventing Language Models From Hiding Their Reasoning
2023/10/27 by Fabien Roger, Roger, Fabien, Ryan Greenblatt +1 · 9 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Stress-Testing Capability Elicitation With Password-Locked Models
2024/05/29 by Ryan Greenblatt, Fabien Roger, Greenblatt, Ryan +5 · 5 citations
Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fault Detection and Control Systems #Industrial Vision Systems and Defect Detection #Machine Learning (cs.LG)
- Benchmarks for Detecting Measurement Tampering
2023/08/29 by Fabien Roger, Ryan Greenblatt, Roger, Fabien +7 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling