vix.ing · top · new · best · stats · spec

Kshitij Sachan

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 101 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  2. Debating with More Persuasive LLMs Leads to More Truthful Answers
    2024/02/09 by Akbir Khan, John Hughes, Khan, Akbir +17 · 57 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Legal Education and Practice Innovations
  3. AI Control: Improving Safety Despite Intentional Subversion
    2023/12/12 by Ryan Greenblatt, Buck Shlegeris, Greenblatt, Ryan +5 · 2 voices · 41 citations
    Computer Science · Engineering · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Safety Systems Engineering in Autonomy #Software Reliability and Analysis Research #cs.LG
  4. Polysemanticity and Capacity in Neural Networks
    2022/10/04 by Adam Scherlis, Kshitij Sachan, Scherlis, Adam +7 · 15 citations
    Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)