Logan Graham
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 118 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
2020/04/15 by Miles Brundage, Brundage, Miles, Shahar Avin +123 · 2 voices · 31 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Law, AI, and Intellectual Property #cs.CY
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
2025/01/31 by Mrinank Sharma, Meg Tong, Sharma, Mrinank +89 · 8 voices · 63 citations
Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes