Evan Hubinger
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
#cs.AI #cs.LG #stat.ML
- Kronfluence: Influence Functions with Eigenvalue-corrected Kronecker-Factored Approximate Curvature
2023/08/07 by Roger Grosse, Juhan Bae, Grosse, Roger +31 · 3 voices · 49 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
- Alignment faking in large language models
2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 63 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Measuring Faithfulness in Chain-of-Thought Reasoning
2023/07/17 by Tamera Lanham, Lanham, Tamera, Anna Chen +59 · 4 voices · 82 citations
Computer Science · #Advanced Graph Neural Networks #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
- Steering Llama 2 via Contrastive Activation Addition
2023/12/09 by Nick Gabrieli, Panickssery, Nina, Gabrieli, Nick +8 · 158 citations
Computer Science · #Interactive and Immersive Displays
- Discovering Language Model Behaviors with Model-Written Evaluations
2022/12/19 by Ethan Perez, Sam Ringer, Perez, Ethan +123 · 120 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
- Risks from Learned Optimization in Advanced Machine Learning Systems
2019/06/05 by Evan Hubinger, Chris van Merwijk, Hubinger, Evan +7 · 34 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning and Algorithms #Reinforcement Learning in Robotics
- Auditing language models for hidden objectives
2025/03/14 by Samuel D. Marks, Samuel Marks, Marks, Samuel +71 · 1 voice · 15 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
2023/07/17 by Ansh Radhakrishnan, Radhakrishnan, Ansh, Karina Nguyen +45 · 5 citations
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
2024/04/25 by Olli Järviniemi, Järviniemi, Olli, Evan Hubinger +1 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering Research #Topic Modeling
- Engineering Monosemanticity in Toy Models
2022/11/16 by Adam S. Jermyn, Nicholas Schiefer, Jermyn, Adam S. +3 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
- An overview of 11 proposals for building safe advanced AI
2020/12/04 by Evan Hubinger, Hubinger, Evan · 3 citations
Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Software Engineering Research
- Agentic Misalignment: How LLMs Could Be Insider Threats
2025/10/05 by Aengus Lynch, Lynch, Aengus, Benjamin Wright +12 · 20 citations
Business, Management and Accounting · Computer Science · #Securities Regulation and Market Practices #Corporate Insolvency and Governance #Cybercrime and Law Enforcement Studies
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
2025/11/03 by Sharan Maiya, Maiya, Sharan, Henning Bartsch +5 · 2 voices · 2 citations
Computer Science · Psychology · #cs.CL #cs.AI #cs.LG
- Conditioning Predictive Models: Risks and Strategies
2023/02/02 by Evan Hubinger, Hubinger, Evan, Adam S. Jermyn +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling