vix.ing · top · new · best · stats · spec

Buck Shlegeris

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 99 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  2. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
    #cs.AI #cs.LG #stat.ML
  3. Alignment faking in large language models
    2024/12/18 by Ryan Greenblatt, Carson Denison, Greenblatt, Ryan +38 · 16 voices · 64 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  4. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
    2022/11/01 by Kevin Wang, Wang, Kevin, Alexandre Variengien +7 · 118 citations
    Computer Science · Materials Science · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Machine Learning in Materials Science
  5. AI Control: Improving Safety Despite Intentional Subversion
    2023/12/12 by Ryan Greenblatt, Greenblatt, Ryan, Buck Shlegeris +5 · 2 voices · 33 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
  6. Supervising strong learners by amplifying weak experts
    2018/10/19 by Paul F. Christiano, Christiano, Paul, Buck Shlegeris +3 · 25 citations
    Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Advanced Bandit Algorithms Research #Machine Learning and Algorithms
  7. Polysemanticity and Capacity in Neural Networks
    2022/10/04 by Adam Scherlis, Kshitij Sachan, Scherlis, Adam +7 · 13 citations
    Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
  8. How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
    2025/04/07 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +5 · 2 voices · 7 citations
    Social Sciences · Computer Science · Psychology · #Ethics and Social Impacts of AI #Adversarial Robustness in Machine Learning #Human-Automation Interaction and Safety
  9. Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
    2024/09/12 by Charlie Griffin, Griffin, Charlie, L. Nicole Thomson +5 · 8 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI
  10. Adversarial Training for High-Stakes Reliability
    2022/05/03 by Daniel M. Ziegler, Ziegler, Daniel M., Seraphina Nix +21 · 3 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Ethics and Social Impacts of AI
  11. Towards evaluations-based safety cases for AI scheming
    2024/10/29 by Mikita Balesni, Balesni, Mikita, Marius Hobbhahn +29 · 7 citations
    Engineering · Health Professions · Decision Sciences · #Safety Systems Engineering in Autonomy #Occupational Health and Safety Research #Risk and Safety Analysis
  12. The Singapore Consensus on Global AI Safety Research Priorities
    2025/06/25 by Yoshua Bengio, Bengio, Yoshua, Tegan Maharaj +171 · 2 voices · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
  13. Language models are better than humans at next-token prediction
    2022/12/21 by Buck Shlegeris, Shlegeris, Buck, Fabien Roger +5 · 3 citations
    Computer Science · #Topic Modeling #Text Readability and Simplification #Natural Language Processing Techniques
  14. Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
    2024/12/17 by Alex Mallen, Mallen, Alex, Charlie Griffin +6 · 5 citations
    Computer Science · Social Sciences · #AI-based Problem Solving and Planning #Ethics and Social Impacts of AI #Adversarial Robustness in Machine Learning
  15. Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
    2024/11/26 by Jiaxin Wen, Vivek Hebbar, Wen, Jiaxin +20 · 4 citations
    Computer Science · #Blockchain Technology Applications and Security
  16. Benchmarks for Detecting Measurement Tampering
    2023/08/29 by Fabien Roger, Roger, Fabien, Ryan Greenblatt +7 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling