vix.ing · top · new · best · stats · spec

Greenblatt, Ryan

  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  2. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
    #cs.AI #cs.LG #stat.ML
  3. Alignment faking in large language models
    2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 63 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  4. AI Control: Improving Safety Despite Intentional Subversion
    2023/12/12 by Ryan Greenblatt, Greenblatt, Ryan, Buck Shlegeris +5 · 2 voices · 32 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
  5. Preventing Language Models From Hiding Their Reasoning
    2023/10/27 by Fabien Roger, Ryan Greenblatt, Roger, Fabien +1 · 9 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  6. Stress-Testing Capability Elicitation With Password-Locked Models
    2024/05/29 by Ryan Greenblatt, Greenblatt, Ryan, Fabien Roger +5 · 5 citations
    Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fault Detection and Control Systems #Industrial Vision Systems and Defect Detection #Machine Learning (cs.LG)
  7. Natural Emergent Misalignment from Reward Hacking in Production RL
    2025/11/23 by MacDiarmid, Monte, Wright, Benjamin, Uesato, Jonathan +19 · 8 citations
    Computer Science · #Topic Modeling #Adversarial Robustness in Machine Learning #Software Engineering Research
  8. Benchmarks for Detecting Measurement Tampering
    2023/08/29 by Fabien Roger, Roger, Fabien, Ryan Greenblatt +7 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling