Mindermann, Sören
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 104 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- The Alignment Problem from a Deep Learning Perspective
2022/08/30 by Ngo, Richard, Chan, Lawrence, Mindermann, Sören · 8 voices · 40 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Alignment faking in large language models
2024/12/18 by Ryan Greenblatt, Carson Denison, Greenblatt, Ryan +38 · 16 voices · 76 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt
2022/06/14 by Sören Mindermann, Mindermann, Sören, Jan Brauner +19 · 28 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Tingchen Fu, Barez, Fazl +39 · 6 voices · 12 citations
Engineering · #Fault Detection and Control Systems
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
2023/09/26 by Lorenzo Pacchiardi, Pacchiardi, Lorenzo, Alex Chan +13 · 18 citations
Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Scientist AI Needs a Government: Structural Governance as the Missing Layer in Non-Agentic AI Safety
2025/02/21 by Bengio, Yoshua, Michael K. Cohen, Damiano Fornasiere +22 · 25 citations
Physics and Astronomy · #Space Science and Extraterrestrial Life
- International AI Safety Report
2025/01/29 by Yoshua Bengio, Sören Mindermann, Bengio, Yoshua +180 · 20 citations
Social Sciences · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Identifying Causal-Effect Inference Failure with Uncertainty-Aware Models
2020/07/01 by Jesson, Andrew, Mindermann, Sören, Shalit, Uri +1 · 3 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Specific versus General Principles for Constitutional AI
2023/10/20 by Sandipan Kundu, Kundu, Sandipan, Yuntao Bai +69 · 5 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- International Scientific Report on the Safety of Advanced AI (Interim Report)
2024/11/05 by Bengio, Yoshua, Mindermann, Sören, Privitera, Daniel +41 · 4 citations
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences
- Active Inverse Reward Design
2018/09/09 by Sören Mindermann, Mindermann, Sören, Rohin Shah +5 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
- In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
2025/04/17 by Ben Bucknall, Saad Siddiqui, Bucknall, Ben +41 · 2 voices · 2 citations
#cs.CY
- Occam's razor is insufficient to infer the preferences of irrational agents
2017/12/15 by Armstrong, Stuart, Mindermann, Sören · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- Bare Minimum Mitigations for Autonomous AI Development
2025/04/21 by Clymer, Joshua, Duan, Isabella, Cundy, Chris +10 · 1 citation
#Computers and Society (cs.CY) #FOS: Computer and information sciences
- International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
2025/10/15 by Yoshua Bengio, Bengio, Yoshua, Stephen Clare +134 · 1 citation
Social Sciences · #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences