Hubinger, Evan
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Korbak, Tomek, Mikita Balesni +79 · 25 voices · 66 citations
#cs.AI #cs.LG #stat.ML
- Kronfluence: Influence Functions with Eigenvalue-corrected Kronecker-Factored Approximate Curvature
2023/08/07 by Roger Grosse, Juhan Bae, Grosse, Roger +31 · 3 voices · 49 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
- Alignment faking in large language models
2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 64 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Measuring Faithfulness in Chain-of-Thought Reasoning
2023/07/17 by Tamera Lanham, Lanham, Tamera, Anna Chen +59 · 4 voices · 83 citations
Computer Science · #Advanced Graph Neural Networks #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
- Steering Llama 2 via Contrastive Activation Addition
2023/12/09 by Panickssery, Nina, Nick Gabrieli, Julian Schulz +8 · 159 citations
Computer Science · #Interactive and Immersive Displays
- Discovering Language Model Behaviors with Model-Written Evaluations
2022/12/19 by Ethan Perez, Perez, Ethan, Sam Ringer +123 · 120 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
- Risks from Learned Optimization in Advanced Machine Learning Systems
2019/06/05 by Evan Hubinger, Chris van Merwijk, Hubinger, Evan +7 · 35 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning and Algorithms #Reinforcement Learning in Robotics
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
2024/06/14 by Denison, Carson, MacDiarmid, Monte, Barez, Fazl +11 · 24 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Auditing language models for hidden objectives
2025/03/14 by Samuel Marks, Samuel D. Marks, Marks, Samuel +71 · 1 voice · 16 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
2023/07/17 by Ansh Radhakrishnan, Radhakrishnan, Ansh, Karina Nguyen +45 · 5 citations
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
2024/04/25 by Olli Järviniemi, Järviniemi, Olli, Evan Hubinger +1 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering Research #Topic Modeling
- Engineering Monosemanticity in Toy Models
2022/11/16 by Adam S. Jermyn, Jermyn, Adam S., Nicholas Schiefer +3 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
- An overview of 11 proposals for building safe advanced AI
2020/12/04 by Evan Hubinger, Hubinger, Evan · 3 citations
Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Software Engineering Research
- Sabotage Evaluations for Frontier Models
2024/10/28 by Benton, Joe, Wagner, Misha, Christiansen, Eric +13 · 6 citations
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Agentic Misalignment: How LLMs Could Be Insider Threats
2025/10/05 by Aengus Lynch, Benjamin Wright, Lynch, Aengus +12 · 20 citations
Business, Management and Accounting · Computer Science · #Securities Regulation and Market Practices #Corporate Insolvency and Governance #Cybercrime and Law Enforcement Studies
- Natural Emergent Misalignment from Reward Hacking in Production RL
2025/11/23 by MacDiarmid, Monte, Wright, Benjamin, Uesato, Jonathan +19 · 8 citations
Computer Science · #Topic Modeling #Adversarial Robustness in Machine Learning #Software Engineering Research
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
2025/11/03 by Sharan Maiya, Maiya, Sharan, Henning Bartsch +5 · 2 voices · 2 citations
Computer Science · Psychology · #cs.CL #cs.AI #cs.LG
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
2025/05/20 by Chiu, Yu Ying, Wang, Zhilin, Maiya, Sharan +4 · 6 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG)
- Conditioning Predictive Models: Risks and Strategies
2023/02/02 by Evan Hubinger, Hubinger, Evan, Adam S. Jermyn +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling