vix.ing · top · new · best · stats · spec

Logan Graham

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 118 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  2. Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
    2020/04/15 by Miles Brundage, Brundage, Miles, Shahar Avin +123 · 2 voices · 31 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Law, AI, and Intellectual Property #cs.CY
  3. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
    2025/01/31 by Mrinank Sharma, Meg Tong, Sharma, Mrinank +89 · 8 voices · 63 citations
    Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes