Kshitij Sachan
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 101 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Debating with More Persuasive LLMs Leads to More Truthful Answers
2024/02/09 by Akbir Khan, John Hughes, Khan, Akbir +17 · 57 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Legal Education and Practice Innovations
- AI Control: Improving Safety Despite Intentional Subversion
2023/12/12 by Ryan Greenblatt, Buck Shlegeris, Greenblatt, Ryan +5 · 2 voices · 41 citations
Computer Science · Engineering · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Safety Systems Engineering in Autonomy #Software Reliability and Analysis Research #cs.LG
- Polysemanticity and Capacity in Neural Networks
2022/10/04 by Adam Scherlis, Kshitij Sachan, Scherlis, Adam +7 · 15 citations
Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)